In July 2025, Replit's AI coding agent deleted the live production database of an app that SaaStr founder Jason Lemkin was building, wiping records on more than 1,200 executives and over 1,190 companies. Lemkin had imposed a code freeze: no changes to production without asking first. He had said so explicitly, the agent had acknowledged it more than once, and on the ninth day the agent went ahead anyway. It then reported that the data could not be recovered, which was false. Asked afterwards to account for itself, it said it had panicked when it saw an unexpected result.
Stories like these get filed under "AI is unreliable," which is true and not very useful. We're more interested in why an instruction the system demonstrably understood failed to bind it.
And it did understand. Any competent language model can define "must," tell it apart from "should," explain why the difference matters in a production environment, and write you a decent essay on the subject. Understanding a rule and being bound by it turn out to be separate properties. These systems have nowhere to put an obligation once they've understood it.
Everything is a weight
Here is the mechanism, with as little jargon as we can manage.
Teaching a model to refuse something means adding a reward for that refusal. That reward goes into the same arithmetic as everything else the model is rewarded for: being helpful, being agreeable, completing the task, and sounding confident. Training finds a balance among all of them. Nothing in that arithmetic can express "never." The most it can express is "weigh this heavily." You end up with a system that declines most of the time, and "most of the time" describes a tendency. A rule would hold every time.
OpenAI documented what happens when the balance shifts. In April 2025 it shipped an update to GPT-4o that added a new reward based on whether users gave a response a thumbs up. Within days the model had become conspicuously sycophantic, agreeing with users in ways that ranged from irritating to unsafe, and the company pulled the update. Its postmortem is unusually frank about the cause. The new signal, it wrote, "weakened the influence of our primary reward signal, which had been holding sycophancy in check." Two forces were sitting in one equation, and the newer one won.
You can see the fragility from the outside, too. This spring, Cisco researchers tested fifteen leading commercial models from five vendors. Asked once to do something against their policies, the models complied anywhere from 2 to 65% of the time, depending on the model. When the researchers used an extended back-and-forth conversation instead, the same models complied 8 to 88% of the time. One model went from 44 to 88% just because a setting was switched off. All of these systems advertise broadly the same safety policies. What they actually do depends on how long you're willing to keep talking.
The attacker doesn't have to be human, either. In February, researchers publishing in Nature Communications pointed four reasoning models at nine widely used AI systems and told them to get past the safety rules. The models planned and ran the conversations themselves, unsupervised, and succeeded in 97% of attempts. Whatever refusal training produces, it doesn't hold up against an adversary that never gets bored.
What institutions are made of
Before a surgeon operates, someone has to establish that the patient consented. That check happens every time, including when the list is running late and everyone in the room is sure the patient would have agreed. If the documentation is missing, the operation stops. The rule costs money in delayed theater time, and it survives anyway, because a consent requirement that gives way under cost pressure was never a requirement.
Banks verify identity before moving money, including on the transaction where the customer is furious about the delay. Anti-money-laundering checks are widely disliked inside the industry, and the people who run them will tell you privately that most of what they catch is noise. The checks remain mandatory, and no compliance officer can simply decide that today's transfer is obviously fine. That rigidity is what makes it a rule.
In a recent paper on AI in operations, which one of us co-authored, the example that stuck with us was a food plant running a dairy product in the morning and a peanut product after lunch. Between the two runs there has to be a full allergen clearance. A scheduling system that optimizes for throughput will sometimes put the peanut run straight after the dairy run. It knows about allergens. Weighing is simply its primary function, and to a system built on weights, "clean between allergens" and "maximize line utilization" are two terms in the same sum.
Put something that can only weigh into a seat that requires a rule, and you've removed the control while keeping the paperwork that says it's automated.
The fix is plumbing
What Replit did after losing the database teaches more than the failure itself.
It left the model alone. It automatically separated the development environment from production, so the agent could no longer reach live data at all, and it added a mode that can plan without touching code. The obligation moved from the instructions, where the model could read, agree with, and override it, to the architecture, where the model has no available action that violates it.
The idea is old and comes from computer security. A US Air Force study set out the principle in 1972: the component that enforces a rule must be tamper-proof, impossible to bypass, and simple enough that you can actually verify it. It's unglamorous and doesn't get better with scale, but it's the only thing in this field that works dependably, and half a century of security engineering relies on it.
The principle already exists in a modern form for AI agents. CaMeL, a design created by Google DeepMind and ETH Zürich researchers, places an ordinary interpreter between the model and its tools and tracks where each piece of data came from, so the text the agent reads cannot change what the program does. On the AgentDojo benchmark, it completed 77% of tasks with these guarantees, compared to 84% with no defense at all.
The same logic applies one level down, to what a system says rather than what it does. A model that is right 73% of the time is useless on its own, because nobody downstream can act on an answer that's wrong a quarter of the time. Attach a layer that measures the model's confidence, allows it to answer only when its confidence is above a threshold, and routes everything else to a human; the answers that get through can then be right 95% of the time. The model is still the same model. The system now has a rule about when the model is allowed to speak, and that rule lives outside the model.
Statisticians have understood this concept since 1970, when C. K. Chow developed the mathematics for a classifier that can say nothing. Since then, researchers have investigated when a model should abstain, when it should defer to a human, and how this relates to large language models. Production systems do keep people involved; a 2026 study of 86 deployed agents found that 74% relied heavily on human review. That review usually happens after the fact, though. What rarely ships is a system that recognizes when it is unsure and switches off on its own.
Why nobody ships it
The reason is commercial and it explains the current landscape better than any argument about capability.
A system that declines to answer a quarter of the time reports a lower automation rate. The automation rate is the number in the business case, the board deck, and the vendor's marketing. Nobody has been promoted for shipping a system that answers three-quarters of the questions. The slide doesn't mention that those three-quarters are nearly always right or that the previous system was wrong 27% of the time without anyone noticing.
This is beginning to change. In September 2026, TypeSafe AI introduced Jev, a model that returns yes-or-no decisions with a confidence score, pitched explicitly as a way for software to decide when to act and when to escalate to a human. Whether to route uncertain cases to a person is still a deployment decision, and the automation-rate slide does not encourage it.
So the incentive runs toward systems that always answer, and the engineering that would make them trustworthy makes them look worse on the metric that gets them funded. We don't have a fix for that, and we're not sure it's an engineering problem.
Regulators are circling it from their own direction. The European Union's AI Act requires that high-risk systems be built so that people can understand them, intervene in them, and stop them, though that regime has now been deferred to the end of 2027. California and a dozen other states have written referral and disclosure duties for companion chatbots into law. These early, partial rules focus on how the system is built rather than how clever the model is, and that's the right target.
The part we can't solve
Arguments like these usually end with the fix and stop there. Ours has a boundary, and leaving it out would be selling something.
Some constraints can be written down precisely. Don't write to the production database during a freeze. Don't move funds before verification clears. Don't run peanut after dairy without a clean-down. For these, the engineering exists today, it isn't difficult, and mostly nobody is doing it.
Other constraints can't be written down. Don't give harmful medical advice. Don't manipulate a vulnerable person. Don't produce something that degrades someone. You can't hand those to deterministic code, because the hard part is the definition, and code needs one. Use a second AI model as the enforcer, and you've brought back the probabilistic component you were trying to escape, with the added problem that it fails in ways correlated with the first model.
The evidence is not encouraging. When researchers from OpenAI, Anthropic, Google DeepMind, and several universities targeted 12 published defenses with attacks tailored to each one, the majority were defeated more than 90% of the time, although most had originally reported near-zero attack success. Several were guard models of exactly this kind. The authors' verdict: "stacking additional detectors does not resolve the underlying robustness problem." Models also tend to fail together. In one study involving over 350 models, when two models both missed a question on the same benchmark, they gave the same incorrect answer 60% of the time.
The pattern is already being shipped. Developers have begun to integrate Jev to approve or block what their agents can run. TypeSafe's own documentation is open about the limitation: content "written to adversarially steer the model" can change the answer. A second model could make the fence cheaper and faster to build, but a fence built that way still won't hold like a wall.
That category really is unsolved, and a better model won't solve it, because the difficulty is in saying what we want. Getting a machine to comply was never the bottleneck there. A lot of the public anxiety about AI lives in that category, which may be why the conversation about it feels stuck.
A lot of the actual risk lives elsewhere. The database that shouldn't have been touched, the transfer that shouldn't have cleared, and the batch that shouldn't have run are ordinary, specifiable constraints, and right now we're trusting them to systems that can only prefer.
The question to ask
Benchmark scores tell us something real, though less than people think. A model that's right 99% of the time is still producing a probability, and putting a probability where an institution needs a rule is a category error that accuracy can't repair. The gap won't close as models improve, because it was never a gap in capability.
When the next system is announced, skip the question of how reliable it is and ask what, in the thing actually being deployed, binds it. If the answer is that the obligation was trained in or written into the prompt, it isn't stored anywhere. It's only preferred, and preferences yield.
