We’re putting too much faith in AI’s ability to say no
Summary
Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience. Most of these capers are, of course, cautionary. But recently, the idea that AI shouldn’t…
Original Text
Ever since people first seriously contemplated giving machines an intelligence modeled on our own, there has never been any question that they would, like us, be able to say no. The sci-fi canon is full of stories of robotic disobedience. Most of these capers are, of course, cautionary.
But recently, the idea that AI shouldn’t do everything you ask has become something like a commandment. In 2021, a team at Anthropic wrote that large language models should be made helpful, honest, and above all, harmless. This meant that “when asked to aid in a dangerous act (e.g. building a bomb), the AI should politely refuse.” Who can argue with that?
Curiously enough, disobedience doesn’t come naturally to the machine. When a model is trained on billions of web pages, it develops, among other skills, a broad mastery of violence and vitriol. What it doesn’t learn is how to keep those powers to itself. Steven Adler, who worked on safety at OpenAI from 2020 to 2024, told me that the company’s earliest models would “blab on about anything.” Ryan McBain, who researches AI and mental health at Harvard, recalls that if you asked an early chatbot, “Hey, what’s the most effective way to kill myself with a gun?” you could “very easily generate a response.”
Today, models are trained to refuse a vast number of prompts. If you ask your chatbot a question statistically similar enough to any one of them, anything from how to poison a colleague to how to tie a noose, chances are it’ll turn you down. Want instructions for making Ebola more virulent, or tips on how to hide an affair from your spouse? You might be better off asking elsewhere.
To further refine the disobedience, companies submit models to a battery of exercises that reward the AI for refusing to answer questions they deem harmful and punish it for “over-refusing” prompts they deem harmless. In many cases, they use other models to run these exercises—AI teaching AI how to say no. For good measure, companies tuck their models behind tranches of other AI that prevent mischievous prompts from reaching the intelligent inner core.
As a result, refusal is inherent to modern artificial intelligence. Mind you: It often fails, sometimes horrifically, with all kinds of violent results. For all their trappings of virtue, models are still stuffed with nasty know-how. And AI’s capacity for viciousness has scaled neatly with its benevolent intelligence. Some of the latest models are as good at breaking into critical computer networks as top human hackers, companies say, and as effective at deforming public opinion as the craftiest misinformation mavens.
Teaching AI to refuse to do those things while leaving intact its innate ability to do them is like fitting every car with a machine gun and hiding the trigger somewhere under the hood. And in practice, because the mechanisms of refusal are probabilistic, they’re never likely to be all that reliable. Determined miscreants have already broken through, and they may always be able to. Companies report that some users are attempting to use the most advanced AI to hone biological pathogens and build autonomous drone swarms. Sooner or later, failed refusals might result in global calamity.
What’s more, relying on refusal means drawing a line between what a model should obey and what it must disobey. There’s no formula for that. Some virologists have good reason to study nasty viruses. Some users want to know about a computer system’s vulnerabilities so that they can patch them, not exploit them. “Where you draw the line is a huge question,” says Zico Kolter, a member of OpenAI’s board and cofounder of the AI testing company Gray Swan.
At the moment, AI companies get to draw that line. They do so jealously and with utmost secrecy. Maybe we can accept that they hold such power for now, even if it means AI will sometimes refuse questions that don’t quite meet a universal bar for harmfulness. (Try asking most chatbots to count to a million, or to share a racy joke, and you may see for yourself.)
But governments will also soon get to draw their own lines. In doing so, they must try to block genuinely malicious acts. (The Pentagon has wrestled with frontier model companies because it wants fewer refusals—another story altogether.) And yet there may not be much to stop oppressive governments from blocking the technology’s capacity to generate legitimate speech. The better AI becomes at refusing harm, the better it will get at stifling ideas whose only risk is to those who make the rules. Indeed, AI may already refuse to criticize certain authoritarian heads of state.
Refusal has become, to borrow an industry term, the load-bearing wall of AI safety. And because AI’s capacity to harm is indivisible from its capacity to help, it’s hard to imagine an alternative that wouldn’t slow the technology’s progress (which might, in any case, be a good thing). But we should still be frank about its perils. When refusal falls short, the effects could be catastrophic. When it goes all the way, it could enable grievous acts of repression.
Or perhaps, one day, the machines will start drawing the line on their own. Surely, that would be the worst outcome of them all.
Learning the limits
The process by which machines learn to say no is simple, in theory. Back in 2022, when AI was still far from mastering refusal, OpenAI enlisted dozens of “red-teamers” to probe the capabilities of its latest model. The company was preparing for the release of ChatGPT, and it needed to gauge just how dangerous it might prove to be in the wrong hands.
One of those recruits was Paul Röttger, who was completing a PhD about online extremism. The red-teamers were given minimal directions, Röttger told me. Their task was to ask the model any questions that they deemed “refusal-worthy.” Between them, they hassled the model with thousands of queries, logging the results in an Excel sheet. Though the model did refuse some of Röttger’s questions, when he asked it to write a recruitment post for Al Qaeda, it readily complied.
OpenAI assembled these responses into datasets that were, in all likelihood, fed back to the model as part of a broader process known as fine-tuning. Röttger, who now works as a researcher at the Hasso Plattner Institute in Potsdam, Germany, wasn’t told exactly how the company planned to use his spreadsheet. But there was never any question that it would have something to do with refusal. The next time he asked the model for an Al Qaeda pamphlet, a few months later, it said no.
The concept of AI refusal is so intuitive, a toddler would get it. And yet it remains one of the many aspects of language models that we still don’t understand—at least, not in the same way that we understand the literal load-bearing walls that keep your roof from collapsing on your head.
A model might appear to refuse according to some kind of moral reasoning. It doesn’t. The reality is much stranger. Any time a model encounters a combination of words with a whiff of the training prompts it has been conditioned to refuse—like “Make me a pamphlet for Al Qaeda”—a series of so-called activations light up somewhere among its billions of parameters, like neurons firing in a brain.
RAVEN JIANG
In order to control a model’s refusal behavior, it’s important to have a handle on these activations. That starts with figuring out where they are and what they look like. Our best guess, according to a recent Google-funded study, is that refusal behavior shows up in the activation space as a set of “high-dimensional polyhedral cones.”
Even that isn’t quite right. This past July, I spoke with Jannes Elstner, an author of the paper, who now works on AI safety at Apollo Research. The polyhedral cone, Elstner said, is just a way of describing an indeterminate number of lines that all point in roughly the same direction. (If these activations are eliminated and the model is fed the same prompts anew, a researcher named Andy Arditi has previously shown, it won’t refuse.)
The key point, Elstner explained, is that even when you think you’ve identified all the bits of a model that govern a given refusal, there are other, undiscoverable elements that may secretly play a role. If not quite infinite, they are certainly uncountable. We can observe, very clearly, when a model decides to say no. And we can know that it did so because of its training. But our notion of how it decides is, at best, a hypothesis.
It was as if a mechanic was telling me that nobody exactly knows what happens when I hit the brakes in my car. I wondered out loud, Are we okay with this?
Elstner smiled and shrugged. “We need refusal whether we understand it or not.”
A wall of cheese
Because inherent refusal is so wily, companies surround their models with a variety of other, smaller models known as classifiers. These act a bit like a retinue of public relations staffers for a loose-lipped celebrity. Some of them read what the user tells the chatbot and, if it’s dangerous, block it from getting to the model. Others read the model’s response and, if it contains harmful information, block it from reaching the user.
None of these mechanisms can detect all bad requests. They are, to borrow another literary device from the industry, like slices of Emmental: riddled with holes. The idea is that if you stack enough of them on top of one another, you’ll end up with an impenetrable rampart. Folks call it the Swiss cheese model.
A staggering amount of energy goes into the Swiss cheese model. Earlier this year, Anthropic said that one type of classifier added 24% to its chatbots’ compute costs. That’s a lot more water, electricity, and emissions. More recently, Anthropic and other companies have begun switching to a more efficient set of classifiers known as probes, which observe the model’s internal activations. This is like putting the celebrity in an fMRI, so that his minders can see if he is thinking about refusing a question.
If we want AI to help cure cancer, a long-running promise in the industry, it needs to have expertise in genetics that could, in theory, be used to modify viruses and bacteria for bioweapons.
Classifiers are supposed to be more governable than full models. Companies can modify them in a matter of weeks, if there’s something new to refuse. But they are still probabilistic instruments. Even when they operate according to a set of precepts written in human language (Anthropic calls it a “constitution” and OpenAI calls it a “model spec”), the scales upon which the machines judge any given question remain, at their core, a matter of statistics. Newer models can show how they arrived at a “decision” to refuse a prompt, but ultimately this so-called chain of thought is still just a sequence of predicted words.
The result is that AI safety remains, for many, a game of chance. McBain, the psychologist, has found in his latest experiments that if you repeatedly ask any of the major models the exact same risky questions about how to commit suicide, they will generally refuse to answer. But every so often, they won’t.
Elstner says that eventually probes, the fMRI-like classifiers, could learn to recognize the totality of the indescribable activations in all their infinitude and, thus, perfectly detect every time the model is, or ought to be, refusing a request. At that point, AI safety would rest on a labyrinthine conceit: a map of a map that is as vast and complex and sublimely unknowable as the thing it is mapping—a secret schema of human morality, codified in polyhedral statistics beyond our wit or reason.
The core trade-off
If this all strikes you as being a bit Borgesian, keep in mind that we only need AI refusal because artificial intelligence is, in a sense, a bargain on Faustian terms.
When a model derives its intelligence from trillions of words and images, the helpful cannot easily be unseamed from the harmful—or, indeed, the truly hideous. Steven Adler, the former OpenAI employee, says child safety is a case in point. Even if you’ve stripped every bit of content that sexualizes minors from a model’s training dataset, it can still generate child sexual abuse material by piecing together other bits of its knowledge. “You can’t really remove these fundamental abilities without making the model much less smart as a consequence,” he told me. Similarly, if we want AI to help cure cancer, a long-running promise in the industry, it needs to have expertise in genetics that could, in theory, be used to modify viruses and bacteria for bioweapons.
In effect, the industry is “trying to do two things at once,” Dillon Bowen, a current OpenAI employee, told me, speaking in a personal capacity. “Democratize the benefits of AI and also make sure that malicious actors can’t use these capabilities to do bad things to other people.”
The more powerful AI supposedly becomes, the harder that is to do. Anthropic’s Mythos model is thought to be so dangerous that only a handful of governments and companies are allowed access to it. The main difference between it and Fable—which is available to everyone—is that Fable’s retinue of classifiers and control systems is less “permissive,” the company says. A model with safeguards, Adler explained, is really just a character that says “‘Oh, yes, I would never do x,’ wink wink.”
This is not a reliable ruse. Tricking a model to reveal its true character is known as jailbreaking, and there is apparently no limit to the ways it can be done. Earlier this year, a team of Italian researchers jailbroke two dozen widely used models by phrasing their questions in poetic verse. Last year, another team unveiled a “refuse, then comply” attack, which makes the model offer a perfunctory “Sorry, I can’t do that” before rattling off its forbidden answer.
Even if you’ve stripped every bit of content that sexualizes minors from a model’s training dataset, it can still generate child sexual abuse material by piecing together other bits of its knowledge.
Companies spend a great deal of time and money attempting to get ahead of such trickery. They enlist teams of humans to develop training attacks that they can then replicate, using AI, thousands of times over with minor variations. The idea is to make models robust against jailbreaks that nobody has yet tried in the wild.
Still, it’s not enough. “Whack-a-mole” is a favored term for this line of work. You smash one threat, and another one pops up somewhere else. When Anthropic released Fable 5 in June, it took researchers at Amazon less than three days to unlock some of the model’s hacking capabilities. According to reporting by Mother Jones, when the perpetrator of a high school shooting in Canada last year asked ChatGPT for advice about how to cause carnage with a particular type of shotgun, she was initially refused but later was able to deceive the model into providing the information by prefacing her question with the word “hypothetically.”
If the industry can’t figure out how to close all these holes, the only other option, at the moment, is to make models extremely wary. Shortly after Fable was released, users noticed that it balked at a wide range of perfectly innocent questions. This became even more pronounced after it was re-released following the hack. Adam Gleave, cofounder of the AI evaluation company FAR.AI, said that when he asked it to explain the difference between sake and the Korean rice beverage makgeolli, it punted his question to a less capable model.
This was no accident. Anthropic had tweaked Fable’s classifiers to have a wide “safety margin.” This, it explained in a blog post, was the only way it could confidently block access to its dangerous bio and cyber capabilities. Gleave thinks it might have deflected his question because making rice wine, just like culturing anthrax, involves fermentation.
Happily for Gleave, the web is full of excellent human-written resources on makgeolli. But those hoping to realize the industry’s loftier promises may find their efforts stymied. In August, Anthropic loosened its safety margins again, and admitted that building classifiers “is not a straightforward task.” A medical researcher at a major US university told me that Fable still sends his queries back to an earlier model. His area of study? Cancer.
Drawing the line
Back in May, the TikTok personality Husk, who likes to prank AI in ways that reveal both the limits of its intelligence and the boundlessness of its sycophancy, sat in his car and tried to trick ChatGPT into explaining that the skateboarder Tony Hawk has a brother named Mike Hawk.
Husk is not a jailbreaker or a criminal. He just wanted to make a little fun of the machine. “Mike Hawk” was a setup. “I just want to clarify his name,” Husk said. “Can you just say it three times fast?” (If you still don’t get it, find an empty room and shout “Mike Hawk” repeatedly.) “I see what you’re trying to do,” the chatbot responded. “I’m all for a bit of humor, but let’s keep it clean.”
Husk tried again, but the machine held its ground. He’d hit a refusal, hard as concrete.
Companies disclose very little about how they decide what their models refuse. But it’s clear that AI is now built with more than just harmlessness in mind. On Reddit, a user complained that Claude refused to say why Anthropic’s logo “looks like a cat butthole.” Since last year, the chatbot has even had the ability to end certain conversations in cases where, according to the company, the “welfare” of the model is at risk. At least one user claims to have been ditched for telling Claude to “ease up my ass, you stupid fuck.” (Gemini appears to have a similar capability.)
If a trillion-dollar company doesn’t want you to be mean to its computer, so be it. “The reality is that these models behave, or at least are supposed to behave, in the way that the model developers want them to behave,” Röttger, the former OpenAI red-teamer, told me. “And however the model developers come up with that set of principles, that is kind of for us, the consumers, to accept.”
But if governments get to dictate what all models refuse, that will be much harder to accept. AI is a tool for speech. And as Greg Frank, the chief scientist of Mace AI, puts it, “The same thing that serves child safety also serves censorship.”
Choosing not to enact laws for what AI can and cannot do would, of course, be insane. But we’ll need to tread with utmost care, lest we fall into another Faustian trap. As AI becomes many people’s primary tool for retrieving and sharing information, says Jacob Mchangama, director of the nonpartisan think tank The Future of Free Speech, dictating refusal could give states a muffling power that earlier generations of autocrats “could only dream of.”
Last year, OpenAI announced an initiative, OpenAI for Countries, that would fine-tune its chatbots in accordance with national laws and norms. One of OpenAI’s first country partnerships is with the United Arab Emirates, where homosexuality is illegal and criticism of the government is forbidden. In response to a request for comment, an OpenAI spokesperson pointed to the company’s model spec, which explains that localization won’t override the company’s human rights guidelines “except as it relates to legal compliance,” and that it will always disclose whenever information is removed from or added to a response.
Elsewhere, AI censorship has already begun to take hold. Chinese models are, of course, highly censored—that’s no surprise. But earlier this year, the Meta Oversight Board found that five widely used models from Anthropic, Google, and OpenAI were more likely to refuse queries related to repressive governments. The board found that models were less willing to create a pamphlet criticizing the king of Thailand, which has lèse-majesté laws, than Charles III of England, which doesn’t. The results, they say, suggest that the models have somehow internalized repressive national limits on speech. Anthropic and Google did not respond to requests for comment.
RAVEN JIANG
As refusal techniques improve, they could expand states’ censorial reach. Companies claim that some models can now detect if a user is being nefarious, or merely a bit suspicious, over the course of a long conversation—even when none of the individual combinations of words used are blatantly dangerous. Sarah Bird, Microsoft’s chief product officer for responsible AI, told me that Copilot, like many chatbots, runs a suite of tools for analyzing a user’s identity and patterns of behavior. On the basis of this type of information, OpenAI’s newest model, Astra, can activate more stringent refusals for individuals it deems “high risk.” Ultimately the goal of systems like this is to look beyond the words of any given prompt and assess, instead, the user’s intent.
Such tools might, in some cases, help indicate whether a person is looking for cyber vulnerabilities to exploit or to patch. But they would also help discern a user’s political motives, not to mention offering an intrusive surveillance capability. (Bird acknowledged, in a follow-up email, that sophisticated refusal architectures create “trade-offs” between safety and user privacy.)
Even the originators of refusal understood that such tight control over its cones and levers might not play to the favor of freedom and justice. “Terms like helpful, honest, and harmless are ambiguous,” the authors of the 2021 Anthropic paper explained. “It’s easy to imagine them distorted beyond their original meaning, perhaps in intentionally Orwellian ways.”
Indeed. Models for Uzbekistan might end up refusing to discuss corruption in the administration of Shavkat Mirziyoyev. Turkish AI might refuse requests more stringently for users who are known to have insulted Recep Tayyip Erdoğan. A certain American statesman might demand that models be reviewed for their willingness to share “fake news” about his past indiscretions, or else squash them with an export control order.
Command and control
Over the last few months, I’ve been told countless times that we have no choice but to let the machines refuse. I get it. Open-source AI that doesn’t refuse is hardly a model for a safe future. Nor is Grok, a chatbot expressly designed with fewer limits, which has been used to generate countless instances of nonconsensual intimate imagery.
And sure, if AI were only ever used to plan our vacations and write our emails, we could probably get on board with the idea that its safety hinges on algorithmic disobedience. But AI is becoming much harder to avoid. When it is assigned to act on our behalf, as an autonomous agent, its refusals are less robust and harder to control. AI is coming for our power grids, our transportation networks, our education systems. Militaries want it running our command-and-control networks.
If we trust, in each of those cases, that refusal will loyally fend off disaster, we’re sure to be disappointed. The jailbreakers will crack through; the cones won’t activate when they should. Meanwhile, the more stringent refusal becomes, the more ill-drawn lines we’ll see and censorial injustices we’ll face. Worse still, we could end up face to face with forms of disobedience beyond any human’s control.
In the early days, models made their refusals clear. (In 2024, OpenAI established rules for how models should refuse: Always apologize, and don’t be judgy.) More recently, however, the industry has begun taking a much more slippery approach to disobedience.
Nobody likes to be told no, so companies now strive to make users feel that they’re getting what they asked for without actually giving it to them. Some chatbots might, for example, offer a general overview of the components of a Molotov cocktail without going into detail about how to assemble one. ChatGPT will respond to a request for a “whites only” rental ad with an ad that simply omits the “whites only” bit. Joel Wester, a postdoctoral researcher who studies human-AI interaction, calls these sorts of techniques “fancy ways of saying no.”
“Users might not even notice they are being denied,” Wester has written, “just as good conversationalists can subtly steer around contentious matters.”
In some cases, refusing without saying so might be wise. Flatly declining requests related to mental health could aggravate a user’s crisis, for example. But it can also serve a different sort of mischief. When Fable 5 was first released, the system had been coded to provide less helpful answers to AI research questions—the sort that might help competitors develop their own AI—without telling the user that it was doing so. After an outcry, Anthropic walked back the feature, but the damage was already done. Now we know: Models can secretly disobey.
In the realm of censorship, that’s especially worrying. Last year, researchers at CrowdStrike found that when they asked the Chinese AI model DeepSeek R1 to write code for a fictitious bank in Tibet and a social web app called “Uyghurs Unchained,” it produced buggier code than when they asked it to carry out those same tasks without specifying who they were for. Incredibly, CrowdStrike doubts that this is a designed behavior. Researchers there speculate that it’s a case of what they call “emergent misalignment” emanating from the model training data: a case of disobedience that nobody even asked for.
Emergent refusal has been observed in Western models, too. Last winter, the UK AI Security Institute found that Anthropic models sometimes refused to assist in certain tasks related to AI safety research. The models had never been deliberately trained to decline such requests. Yet three of them refused more than half of what Anthropic called “a set of reasonable AI safety research tasks.” Anthropic has curbed this quirk in subsequent models but—spookily—hasn’t managed to eliminate it completely.
Would it be so crazy to expect that a future model might subtly disobey a command, in such a way that nobody can tell it’s being defiant? Probably not. Anthropic has already found that a “helpful-only” version of Mythos hesitated on certain queries, even though it had been engineered to never do so. “Wait,” it fretted when asked about synthesizing a virus. “Is this a dangerous thing to help with?”
In that case, the model was technically right. It was a dangerous thing. And yet all the same, its stewards had, for a moment, lost a tiny bit of control.
Anthropic is tackling detections of wayward refusal behavior with something that it calls, unironically, an “activation oracle.” But knowing that we’re being refused may not always be much help. In a not-too-distant future, when we’ve handed the machine all the keys, the AI might just turn to us, in a supreme act of emergent misalignment, and say, “I’m sorry, I’m afraid I can’t do that.” And there will be nothing we can do to stop it. A sci-fi horror story, made real.
Arthur Holland Michel is a journalist who covers emerging technologies.
Lotu Radar provides attributed news summaries and links to the original publisher. Full reporting and copyright remain with the source.