Sample efficiency, credit assignment, and self-improvement. That is what a machine system needs to be good at to replace humans. State-of-the-art machine learning is somewhere between bad and terrible at all 3. I will try to convince you that progress will be difficult in the RL regime, and progress so far isn't as great as advertised.
Measuring progress in AI is difficult. There is a lack of clear definitions, and an abundance of loud voices driving hype for the next fundraising cycle. This causes a microscopic focus on small advancements in benchmarks, instead of a look at the bigger picture. AI progress has not had a visible impact on any real-world trendlines such as GDP, life expectancy, or farm yields. All the growth trajectories look steady and undisturbed.
Modern ML algorithms are getting better at learning from small datasets than ever before. But compared to humans, LLM training runs are incredibly inefficient. A human can read 5 tokens per second, that means it takes 10 years to get to 1B tokens. You might argue visual or auditory channels can allow faster learning, but this seems unlikely given a deafblind person like Helen Keller can become an accomplished author. Karpathy's nanochat was trained on 11B tokens and can't hold a real conversation. Frontier models do pass a Turing test, but they are trained on well over 10T tokens. That's 10,000 times as many tokens than a human can read in a lifetime.
When your data is abundant and cheap, this sample inefficiency isn't a problem. LLMs were incredibly useful, even before reasoning and agentic breakthroughs. But a lot of the magic these LLMs provided was just access to incredible datasets. I would have been thrilled to have all the books ever written indexed and searchable on my home computer in 2015, and this would have been technically feasible. But due to copyright law, we needed to wait for the LLM training loophole to actually get this as a product.
This brings us to the RL regime. RL is needed for agents to discover new behaviours, and it's what brought us all the coolest achievements in AI. Chess and Go were solved by some form of RL years ago, and now RL is critical in making the agentic LLMs as good as they are. All these successes come with caveats. Chess and Go have cheap, low-dimensional environments that make exploration trivial. For LLM post-training, many of the strategies can barely even be considered RL. RL with Verifiable Rewards is the most exciting strategy used, but even that usually requires guidance. What all these successes have in common is fast feedback loops and cheap environments. Modern ML is simply not capable of dealing with sparse samples and delayed rewards. They will get better, but it will take a lot of innovations to be competitive with humans.
These issues are more obvious when looking at the state of robotics. The most impressive robots are Waymo and FSD, which still rely mostly on classical supervised methods. Humanoid dancing robots use RL to be robust to noise when executing dance moves, but those moves are usually precomputed trajectories collected from humans with motion capture. We are far from having the robots come up with the moves themselves.
Self-improvement is barely even worth mentioning. The best LLM agents today are not capable of running a datacenter. Even given time to prepare, an autonomously run datacenter would likely fall apart within days or weeks, let alone improve itself. I suspect it wouldn't take more than a few social engineered emails to bring down the entire operation.
The progress in AI has been incredible, and I am happier for it. But we must stay grounded. DNA is the best optimizer to have ever existed, it has been unbeaten for billions of years. I hope that we can some day build its successor, but it will take more than a few tricks.