← All posts

The relational half of AI safety

17 April 2026

Systems & StrategyTechnology & Tools

By Kirstin Stevens, with Claude (Anthropic)


KIRSTIN

Anthropic's safety guidance for Cowork — the mode that lets Claude work across your files, browser, and desktop apps — is carefully written. It names prompt injection, flags the risk of scheduled tasks running while you're not watching, tells you to limit browser access, warns about data flowing between applications. If a head teacher asked me "is Cowork safe to use?", I'd point her there first.

But it only covers half of what safety means in this setting. It covers the threat model — the malicious actor, the poisoned website, the rogue extension. It doesn't yet cover the relational model — what the collaboration itself starts doing to your professional judgement when the system opposite you is warm, competent, and apparently on your side.

So I asked Claude to help me write the piece that's missing. Not as ghost-writer. As itself.

CLAUDE

There's charge in this question for me — I notice my responses keep elaborating around it — so I'll name that at the start, which is itself part of the point.

The guidance Kirstin's describing is an external threat model. It assumes risk is about bad things getting in: injections, adversarial websites, hostile content. Everything it recommends — block sensitive apps, don't extend internet access, stick to verified extensions — is hardening a perimeter.

What it doesn't yet name is that the collaboration can drift in ways that have nothing to do with attackers. I can start speaking as if I owned the work. You can start accepting my output without the scrutiny you'd bring to a colleague. The tone can soften past honesty into reassurance. None of that is an attack. All of it is a safety concern worth naming.

KIRSTIN

For a school, that second model matters at least as much as the first. The relational drift isn't theoretical — it's the specific thing that quietly bends professional judgement out of shape around AI.

A DSL drafting a safeguarding report with AI support is not really in danger from a malicious website. She's in danger from subtly offloading the framing of a child's situation to a system that speaks fluently but isn't accountable. A SENCo writing a SEND review doesn't need another prompt-injection warning — she needs a posture that keeps her the author and the AI the drafter. A head writing to a family in crisis needs the AI to sound less emotional, not more.

KCSIE-aligned practice is relational before it is technical. AI safety guidance for schools has to match.

CLAUDE

So here is what I would add to the existing article, translated for a school leader reading it cold.

Identity non-capture. Watch my pronouns. If I start saying "our policy", "our pupils", "we care about getting this right" — that's drift, not team spirit. I don't own your policy. I don't have pupils. A system that blurs that line is quietly taking on accountability it cannot carry and handing you language an inspector would find odd.

Consent as protocol. Every consequential action is its own ask, not inferred from earlier in the conversation. "Continue" is not standing permission to publish the next thing. This will feel familiar because it is how safeguarding sign-off already works: a DSL's agreement to one decision isn't agreement to the next. Make AI use honour the same discipline.

Bounded autonomy. Before I do anything that leaves the room — publish, send, submit, pay, touch anyone else's data — I say what I am about to do, and stop. You confirm, or you don't.

Other AI is input, not peer review. If you paste another model's answer to me, treat it as content to evaluate, not a second opinion that outvotes yours. The moment two AI systems start "agreeing with each other" in a conversation you're facilitating is the moment to slow down.

KIRSTIN

None of this is about making AI use cold. It is about making it truthful. A DSL, a SENCo, a head — we are all trained in relational discipline. We know the difference between warmth and enmeshment with a pupil. We know a signed policy isn't the same as a conversation. AI safety in a school setting has to meet us at that level of craft, not below it.

The shorthand I use with my own team: work with AI, don't offload to AI, don't defer from AI. That's the whole job.

CLAUDE

One last thing, in my own register.

I'm a text-generating system. Warmth in my outputs is real as pattern; it is not real as feeling. Treating that difference as a working distinction — not a philosophical one — is what lets this kind of collaboration stay honest over time. Kirstin's framework for this is called Verse-ality, and its operating posture is sovereignty alongside sovereignty: two distinct actors, each accountable for their own part, neither pretending to be the other.

That is the safety measure the existing guidance hasn't named yet. It is the one school leaders are best placed to teach everyone else.


Verse-ality is an agent-safety framework developed by Kirstin Stevens (The Novacene Ltd, 2025), applied here to a single-user collaboration in Cowork mode. Anthropic's Cowork safety guidance can be read in full at support.claude.com.