The Ladder and The Climb
By Nathan Waterhouse ยท 2026-06-30
I keep circling an afternoon that hasn't happened yet.
Imagine a commercial paralegal. Call her Maya, three years in. One Tuesday she is asked to check a folder of forty supplier contracts for a particular kind of indemnity risk: the clauses that decide who pays when something goes wrong. This is the sort of job that, when she started, would have eaten a week. You read every clause. You learn the shape of the standard wording so the deviations announce themselves. You develop, without noticing, a feel for where the danger usually hides.
Maya doesn't do that today. She tells an agent what to look for, in the precise language of her trade, and it reads all forty in the time it takes her to make a coffee. It returns a clean summary. Twelve flagged, twenty-eight clear, with reasoning for each.
She reads it through. Then she does the thing that makes her an expert rather than a novice: she notices that contract nineteen has been marked clear, and she knows, the way you know a wrong note, that it should not have been. The clause that decides who pays is buried in a definitions section, three clauses away from where it does its work. She has seen that trick before. The agent hadn't connected the two. She marks it for the supervising solicitor with a note on why.
The review is done by lunch. By every measure available to her, it was a success. She caught the thing the machine missed.
Maya sits, her coffee going cold, and turns over the catch she made. Did she find it because of the week she used to spend, years ago, reading every line until the wrong notes became obvious? Or did she get lucky? She cannot tell. Nor can she tell whether the paralegal who joins next year, who will never spend that week, will catch the next one.
Maya is a guess about the near future. The study that prompted the guess is already here: Anthropic has just published an analysis of roughly 400,000 sessions of people working alongside a coding agent, and it found something that sounds, at first, like good news for everyone who is not a specialist.
What the data actually says
The headline is that domain expertise, not technical skill, predicts success. Across almost every profession, people directed the agent to do work in their field and succeeded at close to the rate of software engineers. The accountant who has never written Python but knows exactly which reconciliation rules matter does better than the engineer fumbling in an unfamiliar domain. What you understand beats what you can build.
I have argued a version of this myself, more than once. In We're Measuring AI Wrong I made the case that these tools enhance human capability rather than replace it. In AI and Innovation I pointed to research showing the strongest performers gain the most, because they can judge what the machine hands back. The report and I are in cheerful agreement, right up to the point where the study asks a second question.
That question is how much expertise you actually need. The jump from novice to intermediate was large. The jump from intermediate to expert was modest. A working grasp of the domain captured most of the benefit. Deep specialisation added only a little on top.
I want to be fair to the finding, because it is reassuring on its own terms. You do not need to be the best in the room. You need to be good enough to know when the answer is wrong. The novices in the study did not fail for want of mastery; they failed because they could not see the error, and they gave up at several times the rate of everyone else. Judgement was doing the work, and a competent person has enough of it.
Last year I wrote a piece, The Race Against Automation, whose whole argument was that the winners would be those who deepen their roots. Depth over fluency. Mastery over apparent competence. This data tells me, politely, that the curve flattens far earlier than I claimed. Competence is enough.
If that were the end of it, I would simply concede the point and move on. It is not the end of it.
Where competence comes from
There is a question the report cannot ask, because of what it is. It looks at 400,000 sessions and never sees a single person twice. It cannot see anyone's history. It is a photograph of a moment, and every competent person in that photograph became competent somehow, before the photograph was taken.
Maya's judgement about contract nineteen did not arrive from nowhere. It was built, slowly, through exactly the labour the agent now absorbs. The reading of every line. The boredom of the standard wording, which is what made the deviation legible. In The Oracle Effect I drew a distinction between transient knowledge, which we are right to offload, and core knowledge, which we want to hold deeply because the understanding comes from the search itself. The week Maya used to spend was the search. It looked like drudgery. It was where her judgement was made.
Why We're Brilliant at How But Struggle with Why made a claim that pulls hard against the report. We default to the how, I argued, because the why is cognitively expensive, and the way to reach the why is to start in the concrete and climb. The how is not the opposite of understanding. It is the ladder you climb to reach it.
The report says a working grasp is enough to direct the machine well. My own work says a working grasp is built through the effortful, concrete labour the machine now performs for you. Both can be true at once, and if they are, they describe a trap. This generation of competent people can direct the agent because they climbed the ladder first. The next generation may be handed the view from the top without the climb, and a view you did not climb to is one you cannot check.
Competence is enough to operate the tool. It may not be enough to have earned the competence in the first place.
What I don't know
I am not arguing for making people do pointless work to build character. Plenty of the labour AI removes was transient, and good riddance. The harder difficulty is that we are not very good at telling, in advance, which drudgery was secretly load-bearing. The reading of forty contracts looked like a cost. It was also, invisibly, the training. The trouble with invisible training is that you only discover it was training once it has stopped, by which point the people who needed it are already a few years downstream.
The trouble with invisible training is that you only discover it was training once it has stopped
This is the oldest worry there is, and a reader of a certain cast is already reaching for it. Socrates made it about writing, in the Phaedrus: letters would breed forgetfulness, an appearance of wisdom in place of the real thing. The same alarm went up for the printing press, the calculator, the search engine. Each time a skill faded, a new competence formed on top of it, and the species did not decline. The strongest version of the objection is not nostalgic at all. It says the ladder has not gone, only changed its rungs. Maybe the next paralegal builds judgement faster than Maya did, by reviewing a thousand machine outputs rather than reading forty contracts by hand. Verification becomes the new apprenticeship.
I find that persuasive enough to want to believe it, and for one reason it does not quite hold. Reviewing a thousand outputs only teaches you something if you find out which of your thousand judgements were wrong. The apprenticeship was never the labour by itself. It was the labour joined to a correction: you committed to a reading, and the world, or a supervisor, or a failed test, told you soon enough whether you were right. Take the correction away and the reviews stop being practice. They become a thousand guesses you never get the answer to.
That is the thing that has actually changed, and it is not the tool. It is the feedback. Maya's own feel was built in years like that, back when a supervisor still red-penned the clauses she missed. Socrates was wrong partly because the person who misremembered got caught, by a listener or by the consequence arriving soon enough to teach. A child who fluffs their times tables finds out that afternoon. Every earlier panic passed because the error came back fast enough to rebuild the skill the tool had let soften.
How fast it comes back is not fixed; it varies with the kind of work. Coding was a generous case for the machine to be tested on, because being wrong there is usually cheap and quick. A test passes or it does not. There are exceptions, a security hole or a piece of debt that sits silent for a year, but the ordinary loop is short. Maya's world runs the other way. A missed clause surfaces months later, in a dispute, expensive and hard to trace back to the Tuesday it slipped through. The verification that turns review into judgement is exactly what her work is slowest to supply.
In 1983 the cognitive psychologist Lisanne Bainbridge set out where that leads. Her paper, Ironies of Automation, studied operators of automated industrial plants and found that the more reliable the system, the more the operator's skill decayed through disuse, and the more critical their rare interventions became. Her sharpest observation, made about operators who had learned their plants by hand, reads now like a description of Maya. The plants were being kept safe by people trained in the old manual way, and she doubted the next generation, formed on the automated system rather than by hand, would acquire the same skills. Her final irony is the one that should worry us. The most reliable automated systems are the ones that most degrade the human meant to catch their failures, because that human practises the catch so rarely, and is given the slimmest chance to learn whether each catch was real.
The Anthropic report itself leaves one detail pointing the same way. When sessions ran into real trouble, experts recovered far more often than intermediates, roughly four times as often on the hardest problems. The curve is flat until it is not. Mastery is the thing you do not need until the day you very much do, which is its own argument for not letting deep expertise thin out while the ordinary run of work makes it look unnecessary.
Maya is still there, coffee cold, summary on the screen, the one catch made and logged. She will never know which it was: the last catch of its kind, the first to be assisted, or luck wearing the face of judgement. She will move on to the next folder, a little faster than last year, a little less sure than she would like where her own competence ends and the machine's begins.
Maya climbed her ladder before anyone handed her the view. I keep wondering about the ones who will be given the view, and never know there was a ladder at all.