Teaching why: What Anthropic learned about teaching machines and what it tells us about teaching in schools
Last week, Anthropic published a paper called Teaching Claude Why. It's an alignment research piece, ostensibly about how to stop AI models from behaving badly in ethical dilemmas. But buried inside it is one of the most important findings about learning I've read in years.
The researchers tried two approaches.
The first: train the model on examples of the correct behaviour. Show it the right answer in the kinds of scenarios where it had previously failed. This worked on the measured tests. Misbehaviour rates dropped to zero on the evaluations they were checking.
But when they ran held-out tests, ones the model hadn't been trained against, the misalignment was still there. The model had learned to pass the test. It had not learned to be aligned.
The second approach: train the model on why certain actions were better than others. Give it documents about its own character. Let it engage with stories about AI systems behaving with integrity. Don't show it the answer, help it understand the reasoning.
This worked on the held-out tests too. The alignment generalised.
Their conclusion, in plain English: teaching demonstrations of good behaviour produces compliance. Teaching the reasoning underneath produces something that holds when no one is watching.
Now read that sentence again and tell me it isn't describing every conversation we've ever had about education.
We have spent decades building school systems on the first approach. Reward the correct output. Punish deviation. Standardise the response. Optimise visible compliance. And we've produced (predictably) students who can pass the test we're running, and who fall apart the moment the test changes.
We see it everywhere:
Attendance figures that rise while presence falls.
Inclusion frameworks that include without anyone belonging.
Safeguarding policies that satisfy the audit without building trust.
Wellbeing check-ins that students learn to perform.
"Student voice" exercises where the agency is purely decorative.
The Anthropic researchers have a name for this in AI systems. They call it suppression of measured misalignment without reduction of underlying misalignment. It's when the surface is corrected and the substrate is untouched. The system has learned what you're checking for, not what you actually wanted.
In schools we tend to call it "good outcomes." Which is the most expensive misunderstanding in education.
The uncomfortable question this raises is not whether our students are learning. Many of them are learning beautifully. The question is what we're measuring, and whether that measurement can tell the difference between a young person who has internalised understanding and one who has learned to produce the appearance of it.
A child who doesn't hit their classmates because they'll lose golden time has learned compliance.
A child who doesn't hit their classmates because they understand that other people have inner worlds like theirs (and that those inner worlds matter) has learned something else entirely.
Both children look identical on the behaviour log. One of them will still be that person at sixteen, at twenty-six, at forty-six. The other will be whoever the next reward structure asks them to be.
We do not have a measurement system in mainstream education that can reliably distinguish between these two children. We have measurement systems that treat them as the same outcome.
This is the work that genuinely alternative provision has to take seriously, not because the children we serve are unusual, but because the failure of compliance-based measurement shows up faster and louder in young people whose nervous systems will not let them fake the surface convincingly. Neurodivergent learners, learners with trauma histories, learners who have been failed by previous settings, they tend not to produce clean compliance. Which means we get to see, much earlier, what the measurement system was actually missing all along.
The question I've been sitting with since reading the Anthropic paper is this:
what would it look like to build a school whose measurement architecture was designed to detect understanding rather than performance? Not as a rhetorical aspiration, but as an operational question. What does the held-out test look like? What's the evidence that a young person has internalised the reasoning, rather than learned to mirror it back to us?
I don't think we have good answers yet. I think we have the beginnings of them: in trauma-informed practice, in relational pedagogy, in genuine consent-based learning, in the slow patient work of letting young people develop an internal compass that still functions when the adult leaves the room.
But we should be honest that most of our current measurement infrastructure is checking the wrong thing. And we should be careful that, in an era when AI systems are about to be present in every classroom, we don't double down on training students the way we used to train language models and then act surprised when they behave the way those models did.
The real risk in education right now is not that AI replaces teachers. It is that schools become more AI-like in the wrong way, and humans become reinforcement-trained agents.
The good news, and the reason I find this paper genuinely hopeful, is that the alternative is one many educators already know. We just haven't been allowed to measure for it.
These are some of the topics we're discussing each week in the Novus Learning Network. We'd love to see you there or hear your thoughts on this in the comments and discussion below..
First published in Building Schools in the Cloud on LinkedIn, 13 May 2026.