That's exactly what Anthropic said was going to happen!
Their big bet is that models are going to keep getting sharply better, not that they're going to quickly reach a plateau of quality that they can then defend.
They will get sharply better in tasks with verifiable domains...
math and coding
Gradually the labs will start engineering verifiable sandboxes for wider domains like videogames
This strategy will hit a plateau in about 18 months and then we're back to diminishing returns and incremental progress along other dimensions (like accelerated inference using ASICs)
RL can do behavior cloning, but really needs good simulations or verifiable environments to get to superhuman levels. That currently exists for math, coding, and a lot of videogames. Soon there will be good enough simulations for robotics.
There's a lot of domains where that simply isn't the case (like bio)
You get much better supervised data in bio/chem though. These data companies have people working on exactly that.
While it's not going to give you an "alphago" effect, it is still enough to work at human levels, augmented with the general knowledge of an LLM, together making it super-human.
Yes. But there is also no other choice for people in these professions. The underlying job has been automated already. What's left is automating the last leg.
If you consider a 5-year outlook, it is also a very temporary job unless you're like a specialist neurosurgeon or something, as one of the examples in that article shows:
> The on-again, off-again nature of the work is not just the result of company culture; it stems from the cadence of AI development itself. People across the industry described the pattern. A model builder, like OpenAI or Anthropic, discovers that its model is weak on chemistry, so it pays a data vendor like Mercor or Scale AI to find chemists to make data. The chemists do tasks until there is a sufficient quantity for a batch to go back to the lab, and the job is paused until the lab sees how the data affects the model. Maybe the lab moves forward, but this time, it’s asking for a slightly different type of data. When the job resumes, the vendor discovers the new instructions make the tasks take longer, which means the cost estimate the vendor gave the lab is now wrong, which means the vendor cuts pay or tries to get workers to move faster. The new batch of data is delivered, and the job is paused once more. Maybe the lab changes its data requirements again, discovers it has enough data, and ends the project or decides to go with another vendor entirely. Maybe now the lab wants only organic chemists and everyone without the relevant background gets taken off the project. Next, it’s biology data that’s in demand, or architectural sketches, or K–12 syllabus design.
Their big bet is that models are going to keep getting sharply better, not that they're going to quickly reach a plateau of quality that they can then defend.