NotesScienceEdition 14 min read

Build the Conscience In

We test AI the way we test a fortress — attack it, then defend it. A few years ago Adam Russell and I argued that's only two-thirds of the job, and named the missing third.

The first page of the paper The Promise and Peril of Artificial Intelligence — Violet Teaming Offers a Balanced Path Forward, by Alexander Titus and Adam Russell

Security has team colors.

If you want to know whether a system holds, you hire a red team to attack it — to think like an adversary, find the cracks, break in the way someone hostile actually would. Then a blue team builds the defenses: the monitoring, the patches, the walls. Put them in the same room and you get purple teaming, attack and defense sharpening each other.

It’s a good system. We borrowed it from cybersecurity to pressure-test AI, and it works about as well on a language model as it does on a network. Break it, then defend it.

A few years ago, Adam Russell and I kept hitting the same edge of it.

Red and blue tell you whether a system can be broken and whether it can be guarded. Neither one asks what the thing does to the people on the other side of the screen once it works exactly as designed. The bias that ships in a working model. The manipulation a perfectly functional system makes cheap. The person a flawless automated decision quietly denies. None of that is a bug you can red-team your way to. It’s the system doing its job.

So we proposed a third color.

We called it violet teaming. It keeps the red and the blue — you still attack, you still defend — but it adds a commitment the other two don’t carry: build the social benefit in, at design time, as a property of the system rather than a fence you put around it afterward. Responsible by design, not by apology.

That cashes out in two moves. The first is technical, and a little counterintuitive: use the technology itself as part of the safeguard.

Most of our instinct for AI safety is external. Audit it, regulate it, convene a review board, write the law. All necessary. But violet teaming says the capability that creates the risk can also be turned against it. In biology, that looks like running AI screening on a generated sequence during inference — so a model capable of designing something dangerous also checks its own output before it ever reaches a synthesizer. You inoculate the technology with a dose of itself. DARPA has funded work in this direction; researchers have built models that self-destruct rather than be quietly retrained toward harm. Safety becomes something the system carries, not just something we point at it from outside.

The second move is harder, because it isn’t technical at all.

The risks and rewards of AI don’t actually live inside the model. They live in the world where the model meets people — psychology, institutions, incentives, law, the messy sociology of how a tool gets used once it’s loose. We called that the sociotechnical level, the place where the hardware and software meet the humans and reshape each other. You cannot get to safety by staying at the purely technical layer, because the harm isn’t purely technical. Which means the room can’t only hold engineers.

That’s the part the name is pointing at.

Red and blue teaming test the tool. But a tool, by itself, has no conscience. It executes what it was built to do, faithfully, including the damage. The conscience has to be built into the system around it — the design choices, the humans kept in the loop, the values encoded before the thing ships and the people accountable after. Violet teaming is the discipline of putting a conscience where the tool doesn’t have one, and can’t.

And it doesn’t finish. There’s no moment where you shrink-wrap an AI and declare it safe, because the model keeps learning, the world keeps moving, and every tool we make remakes the world it’s used in. We are, as we wrote it, tool-creating apes — we change ourselves by what we build. Safety by design isn’t a milestone you clear. It’s a practice you don’t get to stop.

I’ve kept coming back to this. When I testified to a Senate forum on AI, violet teaming was the mechanism I offered for acting on the risks of AI in the life sciences without strangling the science that makes it worth doing. When the National Academies studied AI and biosecurity, the report cited it. Underneath both is the same refusal of the easy positions: you don’t get safety by banning the technology, and you don’t get it by trusting it. You get it by building responsibility into the technology and the people around it, deliberately, and then again, and again.

It is also, if I’m honest, exactly what my novels are about, run from the other side. The fiction keeps staging the world where nobody built the conscience in — where a capability arrives, works, and does its damage before anyone reckoned with the social layer it landed in. A tool has no conscience. I’ve used that line about a book. It isn’t a complaint about the tool. It’s a job description for us.

Adam and I ended the paper on a sentence I still believe. With conscience and wisdom, the extraordinary capabilities of AI can enrich humanity. But without adequate precaution, the risks could prove catastrophic.

The capabilities were never the part in doubt. The conscience is the part we have to build, and keep building — because there is no finish line, and the tool is never going to grow one on its own.

Titus

Read the novels

Echoes of Tomorrow is complete and available now. Monarch begins withZero Billionaires on September 29, with a prequel trilogy in January 2027.

All Notes