There's a well-known fear in AI research called model collapse. Train models on the output of models, for enough generations, and the distribution degrades - the tails thin out, the variance dies, and eventually you get a confident, fluent, subtly wrong machine that has drifted away from reality with no way to detect the drift.
The Other Collapse
The industry takes this seriously. There are papers, mitigations, expensive fresh-data acquisition programmes, entire vendor categories.
Meanwhile the identical process is running on people, at full speed, with no papers, no mitigations, and no one's name on the problem.
How Expertise Was Actually Manufactured
We tell a flattering story about expertise: study hard, learn from the best, absorb the knowledge. It's mostly false, and every field that has measured it knows it's false.
Expertise is built almost entirely from feedback on your own errors. Not exposure to correct answers - production of wrong ones, followed by correction. This is one of the most replicated findings in skill acquisition research, across motor learning, medicine, aviation, chess, and language. Watching a correct solution produces a strong, immediate feeling of understanding and remarkably little durable skill. Generating a wrong solution and getting it corrected feels terrible and builds the thing.
Which means the apprenticeship system was not inefficient. It was expensively correct.
The junior lawyer buried in document review, the resident on their fourth night shift, the analyst rebuilding a model at 2am, the junior engineer bisecting a bug for six hours - this looked like cheap labour doing grunt work. It was actually a training loop disguised as a cost centre. The organisation paid for slow, wasteful, error-strewn junior output, and what it received in exchange, ten years later, was a senior who had judgment. The grunt work was the substrate the correction loop ran on.
Now here is the part everyone gets wrong. AI did not take the senior's job. It took the junior's tasks - which is to say, it took the substrate. The loop still exists. It just has nothing to run on.
The senior stays excellent. Their judgment was built in the old regime and it doesn't degrade. The junior gets a first draft that is better than anything they could produce, corrects it lightly, ships it, and - this is the critical part - never generates the productive error. They accumulate exposure to correct answers, which is the single least effective input for building judgment that we know of.
"We built a machine that removes the productive error from the workflow. We call it a productivity tool."
Why It's Invisible Until It Isn't
This is the property that makes it dangerous: it has a delay fuse measured in careers.
Expertise has a half-life of about one working life. The seniors trained by reality are in place, they're good, and they'll be good for another fifteen or twenty years. Every metric you'd look at - output quality, error rates, throughput - improves during this period, because you have excellent human judgment reviewing excellent machine output. It is the best of both regimes and it looks like a triumph.
The pipeline behind them is empty, and nothing measures a pipeline until you draw from it.
So the curve is not a decline. It's a plateau followed by a cliff, and the cliff arrives roughly one generation after the cause, at which point the cause is twenty years in the past, unattributable, and universally described as a mysterious skills shortage.
We are, right now, inside the plateau. And the cohort currently holding senior judgment is the last one that was trained by unmediated reality rather than by a model's output. That's not a rhetorical flourish - it's a demographic fact with a date on it.
The Inversion: The Competitive Value of a Worse Tool
Here's where it gets interesting instead of just gloomy.
If judgment is manufactured by unaided error, and everyone removes unaided error from their workflow, then judgment becomes scarce - and anything scarce becomes valuable. Which means: deliberately withholding AI from some fraction of junior work is a capital investment, and the organisations that make it will own the entire supply of senior talent in fifteen years.
This sounds insane in 2026. It sounds like refusing to use calculators. It will sound obvious in 2040, in the same way that "maybe don't outsource your entire manufacturing base and keep zero fabs" sounds obvious now and sounded insane in 1995.
The mechanism is identical, by the way. Offshoring didn't fail because the offshored work was done badly. It was done well and cheaply. It failed - where it failed - because the capability to do the work atrophied domestically, invisibly, over twenty years, and turned out to be unbuyable later at any price. Capability is a stock, not a flow. You notice you spent it long after it's gone.
An organisation that runs a deliberate "unaided track" - some percentage of real work done by juniors without assistance, slower and worse, with senior correction - is paying a visible efficiency cost today to hold a capability stock that will be unpurchasable later. That is exactly what a strategic reserve is, and exactly why nobody with quarterly targets will build one.
The Twist That Should Worry AI Labs Specifically
Now the part I think is genuinely under-argued, and it points at the labs rather than at their customers.
Frontier progress depends on the existence of humans who can judge frontier output.
Every improvement pathway - preference data, evaluation design, red-teaming, benchmark construction, verifying that a claimed scientific result is actually a result - bottoms out in a human who is good enough to tell whether the output is good. In easy domains this doesn't bind; you can check with code, or a test suite, or ground truth. In hard domains, the only available grader is an expert human.
If expert humans stop being manufactured, the grading population fossilises. And a model cannot be reliably pushed past the level of its best available judge in domains where no cheaper ground truth exists. Not because of some deep theoretical limit - simply because you can no longer tell which of two outputs is better, so you can't select for the better one, so improvement in that domain goes flat.
So the ceiling on superhuman performance in unverifiable domains is partly set by the supply of humans capable of recognising it. Which gives AI labs a purely selfish, non-altruistic reason to fund human expertise development - they are consuming a resource they are simultaneously interrupting the production of.
And there's already evidence this is being felt: the sharp rise in what labs pay for genuine domain-expert annotation. That's not charity, it's a scarcity signal. The market for expert judgment is repricing before anyone has named why.
What to Actually Do: Invert the Loop
The practical version fits in one line, and it's a reversal of how essentially every AI workflow is currently built.
"Today: AI does the work, human reviews it. Instead: human does the work, AI reviews it."
Same tools. Same tokens. Completely different training gradient. In the first arrangement, the human is exposed to correct answers and produces no error - the worst learning configuration known. In the second, the human generates real attempts, and gets fast, patient, unlimited, non-judgmental correction - which is, unambiguously, the best learning configuration ever constructed. A tireless expert reviewer available to every junior on earth at 2am is genuinely the greatest apprenticeship technology ever invented.
We got handed the best teaching instrument in history and immediately deployed it as a substitute for the student.
Concretely, for a team: pick a fraction - 20% is a defensible starting number - of real, consequential junior work that is done unaided. Run the AI as the grader on that work, not the producer; have it critique, not draft. Protect it in the budget explicitly, as a training line item, so the first cost-cutting cycle doesn't quietly delete it. It will be the first thing cut otherwise, because it looks exactly like waste. It is waste, in the same way a fire drill is waste.
And track one number nobody tracks: time-to-independent-judgment for new hires. If it's lengthening while output quality is rising, you have found the plateau.
Predictions, So This Can Be Wrong
Through 2030, junior hiring in knowledge professions stays structurally depressed (already visible), followed by a senior wage spike that arrives too late to be causally attributed to it.
Through 2030, rates paid to genuine domain-expert annotators keep rising faster than general wages.
By 2032, at least one high-consequence field - medicine, aviation, structural engineering, or law - introduces a formal AI-free training requirement, and frames it as a competence standard rather than a nostalgia measure.
The kill shot: if we find a way to build judgment from reviewing correct output rather than producing wrong output, the entire argument collapses. That would be a genuinely enormous discovery and I would love to be wrong this specific way.
The One Line
Model collapse is what happens when a system learns only from its own output and loses contact with reality. We solved it for the models. We are running it, unmitigated and unmeasured, on ourselves - and the results won't be visible until the last people who learned from reality have retired.
Part three of three. The triptych, in one sentence: AI made intelligence cheap, and in doing so made everything adjacent to intelligence - accountability, depth, and the manufacture of judgment - the only scarce things left.