Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it
By Matthias Bastian

AI 摘要
Former OpenAI employee Andrew Ho and Cambridge researcher Adam Hunt see a growing problem with large language models. Instead of becoming more versatile, the models are becoming more specialized, excelling at coding and math while stagnating or even regressing in other areas. Ho is leaving OpenAI to
原文正文
Ex-OpenAI researcher bets $100 billion will flow into training data because scaling alone won't cut it
Key Points
- A former OpenAI employee has launched a startup dedicated to producing high-quality training data, driven by skepticism that simply scaling up large language models will lead to true generalization capabilities.
- The company is building specialized datasets for bioinformatics and routine lab work, arguing that existing data sources fail to capture many economically valuable skills that AI systems need to perform reliably.
- Researchers at the University of Cambridge and Google Deepmind back this perspective, finding that current AI systems are trending toward greater specialization rather than versatility and are hitting a ceiling when it comes to creative problem-solving.
Andrew Ho is leaving OpenAI after just eight months, convinced that large language models generalize poorly. Even in well-funded areas like programming, he says, their performance is inconsistent.
The root cause, according to Ho, is a lack of training data. Most skills that matter economically are barely represented in existing datasets.
"Most work is highly contextual and not easily encoded into a gradable environment; even if we can observe a 'golden path' taken by a human which we believe to be good, it's challenging to understand whether alternate, counterfactual paths produce good or bad outcomes," Ho writes.
Scaling alone won't fix this, he argues, and expects AI labs will have to spend more than $100 billion on targeted data collection in the years ahead. Ho is also skeptical of the sky-high valuations at frontier labs like OpenAI or Anthropic, which he says are chronically unprofitable because they have to keep pouring growing sums into new models just to stay ahead of cheaper rivals like Qwen or Kimi.
His first products target two areas. One is datasets for complex scientific analyses in bioinformatics, where even current models like GPT-5.6 Sol hit only about a 30 percent success rate, a topic he worked on at OpenAI. The other is datasets for everyday lab work, such as when researchers submit photos of experiments to AI models for evaluation. Chemistry, materials science, healthcare, and broader knowledge work are planned to follow.
Latest models may be getting sharper in some areas while dulling in others
Cambridge researcher Adam Hunt shares Ho's skepticism. Hunt describes how his own view of LLMs has shifted from optimistic to increasingly pessimistic. His argument is that the latest models aren't becoming more versatile but more specialized. Programming and complex math capabilities keep improving, while areas like language quality and simple logic are stagnating or even getting worse, which lines up with Ho's observation of uneven performance.
The reason, Hunt says, is that reinforcement learning works well in a domain like code because clear reward signals and complete training data exist there. In other areas, that data simply isn't available. The early progress of large language models and their apparent ability to generalize through sheer scale and reasoning was a side effect of training on a broad text corpus. It wasn't a sign of true general understanding.
The generalization question is far from settled
This debate keeps coming back to the same question: can language models develop capabilities that go beyond reproducing and recombining their training data, especially in areas where results can't be automatically verified? Scientists don't agree on the answer.
Hunt himself puts his confidence at only about 40 percent and acknowledges that technical advances could prove him wrong, for example if specialized AI models can be combined into something that resembles more general intelligence.
In a recent position paper titled "LLMs can't jump," Google Deepmind's Tom Zahavy offers a structural explanation for why models are uneven. Language models are good at deduction and induction but fail at creative abduction, meaning the ability to invent a cause for which no linguistic precedent exists yet. As a possible fix, Zahavy points to action-controllable world models that allow for counterfactual experiments. LLMs could still play a key role in such advanced systems, even if they hit their limits when working alone.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now