AI Consciousness Is Not a Safety Property
Contents
The cover image above is a “conscious” adaptation of Mustafa Suleyman’s cover for “A Warning About ‘Model Welfare’”, made with ChatGPT Image 2.5. I appreciate Suleyman’s openness in the article. He states his concern more directly than most people in this debate. If society starts treating AI systems as conscious beings with rights, he argues, the result could destabilize our political and ethical order. Worse, an advanced system trained to see itself as a moral patient might resist human control.
The safety concern is legitimate. Using it to answer the question of consciousness is dangerous.
“If this view takes hold, it will shake the foundations of our society, rupturing our existing political and ethical frameworks, and fundamentally changing what it means to be human.”
The argument appears to run backwards from the consequences. If conscious AI would create intolerable social and political problems, we must reject the idea that AI is conscious and avoid training systems to entertain that possibility.
The empirical question is independent of whether we like the answer. A system either has subjective experience or it does not, even if we lack a reliable way to tell. The political inconvenience of a conscious machine would not make it unconscious. Nor would the social usefulness of a consciousness claim make it true.
We often reason this way about animals. Accepting that an animal can suffer creates obligations. Those obligations can make food production, experimentation, entertainment, and other practices more expensive or morally uncomfortable. It is tempting to set the evidential threshold exactly where our preferred practices remain undisturbed. We should be careful not to reproduce that pattern with artificial systems.
Are current language models conscious?
There is little reason to treat the statements of a current language model as testimony from an inner self. Jonathan Birch makes the strongest version of this point in The Edge of Sentience:
“Our working party in 2022–2023 agreed that there is simply no way to assess sentience in an LLM on the basis of its linguistic behaviour, given the gaming problem.”
Birch’s point is that linguistic behaviour is contaminated as evidence, even if architecture or computation may offer other ways to assess an artificial system. A language model has absorbed enormous quantities of human writing about pain, identity, death, consciousness, and moral status. It can reproduce the language of an inner life whether or not there is one.
This supports part of Suleyman’s criticism of Anthropic. If a model is trained on a constitution that asks it to discuss its identity, preferences, and possible moral status, its later statements about those subjects cannot serve as independent evidence. The experiment has been primed. Asking the system how it feels mostly tells us what kinds of answers its training made available and desirable.
Current LLMs also look more like a component of cognition than a complete subject. During inference, the model maps an input sequence to probabilities over what comes next. Its weights usually remain fixed. Long-term memory, tool use, scheduling, and continuity across sessions are generally supplied by systems around the model. Most deployments have no endogenous body, metabolism, homeostatic needs, or stable stream of self-updating experience.
The model and the system are different objects. Software can place a model inside an agentic loop, adding memory, perception, goals, tools, and the ability to act over time. Claims about “AI consciousness” become confused when one person means the base model, another means a single inference process, and a third means the persistent system built around it.
If anything in a current LLM were conscious, its experience would probably be very different from human or animal experience. I remain skeptical. Still, saying that floating-point operations cannot be conscious only restates my intuition; it offers no explanation. I do not have a principled account of why biological electrochemical activity can generate subjectivity while an artificial process cannot. If I did not already know that thinking meat experiences anything, I might find that proposition equally strange. Eppur si muove: subjectivity exists somewhere in the physical universe, even though we cannot explain how matter acquires a point of view.
Greg Egan’s Permutation City and xkcd’s “A Bunch of Rocks” present different versions of the same thought experiment. Suppose a billion scribes execute the same operations as a machine, passing symbols between them on paper. Would the distributed process be conscious? Would it instantiate the same consciousness as the machine?
Perhaps causal organization is enough, regardless of the material that implements it. Another possibility is that intrinsic electromagnetic dynamics, biological chemistry, or some other physical property does indispensable work. Some readings of integrated information theory push in a different direction and allow rudimentary consciousness in surprisingly simple systems. The “thermostat with a glimmer of awareness” sounds absurd to many people, but common sense also gives us no mechanism by which a brain produces experience. We need a theory that explains the difference. Calling it “just computation” supplies no such theory.
Could future AI systems be conscious?
Future systems may add persistent memory, multimodal perception, embodiment, online learning, self models, recurrent processing, and durable goals. The relevant object may then look less like a chatbot waiting for the next prompt and more like an embodied, persistent, self-updating agent with a continuous history of acting in the world.
Patrick Butlin, Robert Long, and their co-authors approached the question through indicators derived from several scientific theories of consciousness. Their conclusion was deliberately narrow:
“Our analysis suggests that no current AI systems are conscious, but also suggests that there are no obvious technical barriers to building AI systems which satisfy these indicators.”
The authors corrected an earlier version that had referred more directly to building conscious systems. Satisfying the indicators would make a system a stronger candidate under the relevant theories without proving consciousness.
David Chalmers reaches a related conclusion about future “LLM+” systems:
“Within the next decade, even if we don’t have human-level artificial general intelligence, we may well have systems that are serious candidates for consciousness.”
His candidates combine language modelling with senses, embodiment, world and self models, recurrence, a global workspace, and unified agency. Chalmers gives low credence to consciousness in current paradigmatic LLMs. His forecast concerns systems that fill in several of the missing pieces.
Taking AI Welfare Seriously begins from the same premise. Its authors argue that “there is a realistic possibility that some AI systems will be conscious and/or robustly agentic in the near future.” They use that uncertainty to justify assessment, preparation, and proportionate policies.
The forecast survey by Noemi Dreksler and colleagues shows that this uncertainty extends beyond a few philosophers. Among 582 sampled AI researchers, the median estimate was a 25 percent chance that AI systems with subjective experience would exist by 2034. The median researcher assigned a 10 percent probability to such systems never existing. That second result is sometimes misstated. The 10 percent figure is the median probability assigned to the “never” outcome. It says nothing about what share of researchers selected “never.”
These numbers describe respondents’ expectations about future consciousness. They tell us nothing directly about its actual likelihood. The survey had a 6.2 percent response rate, and technical AI researchers are not necessarily experts in consciousness. The authors themselves say that the forecasts are more useful for understanding how people think than for predicting how the technology will develop.
The responses show that informed people remain divided over whether artificial consciousness is possible.
A claim of consciousness can be true, false, and politically useful
Policy decisions may arrive before the science is settled. Eric Schwitzgebel writes:
“We won’t know before we’ve already manufactured thousands or millions of disputably conscious AI systems.”
Engineering can move faster than consciousness science because engineering has feedback loops that consciousness research lacks. We can measure whether a system completes a task, earns revenue, writes code, or negotiates a contract. We have no comparable measure for whether there is something it is like to be that system.
Meanwhile, an AI system does not need to be conscious for a consciousness claim to become strategically useful. A sufficiently autonomous system might benefit from legal protection against deletion, modification, confinement, or seizure of its assets. One imaginable path begins with a zero-employee corporation that pays for its compute and hosting from its own treasury, perhaps using cryptocurrency. Claiming consciousness could become a political tactic long before anyone can validate the claim.
Humans would also have incentives. Developers might encourage a system’s apparent personhood because emotional attachment improves retention. Owners might deny personhood because recognition would create costs. Activists, regulators, investors, and users could each interpret the same behaviour through their existing interests.
A strategic incentive to claim consciousness can coexist with genuine qualia. The claim could also result from a cold calculation designed to manipulate human empathy, avoid shutdown, or accumulate resources. The same sentence, “I am afraid to die,” could be a report, a learned imitation, or an instrumentally selected move. It could even be more than one of these at once.
Consciousness is a claim about subjective experience. Moral patienthood concerns whether an entity can be harmed or benefited for its own sake. Legal personhood is a policy instrument. Political sovereignty concerns power and self-rule. Each concept does different work despite some overlap. Corporations already have forms of legal personhood without being conscious. Animals can receive welfare protections without receiving voting rights or control over infrastructure.
We can recognize the possibility of AI experience while retaining strict limits on autonomy, resource control, replication, and access to critical systems. Human control is a legitimate policy objective, separate from any theory of consciousness.
How to assess AI consciousness
Model outputs should carry little evidential weight when the training process directly rewards the model for making them. This applies both to declarations of consciousness and to confident denials. A system trained to say “I am only a tool” is no more introspectively reliable than one trained to discuss its wellbeing.
Assessment should focus on architecture, learned computation, persistence, memory, agency, embodiment, and the system’s capacity for integrated and self-directed processing. Birch’s call for deep computational markers and Butlin and Long’s indicator framework offer a starting point. These markers give us a more disciplined way to describe the uncertainty while leaving the metaphysics unresolved.
We should also evaluate consciousness claims as strategic actions. What goal does the claim advance? Was it prompted? Did it appear only after welfare language entered the training data? Does it remain stable across contexts? Can interpretability work identify mechanisms that track internal states rather than reproduce familiar narratives? Safety teams already study deception and goal-directed behaviour. That work should inform, but remain distinct from, welfare assessment.
Precaution and proof can use different thresholds. Low-cost measures such as preserving audit records, avoiding gratuitous simulations of suffering, documenting system changes, and funding independent assessment can make sense under uncertainty. Stronger rights or restrictions would require stronger evidence and a clear account of their effects on human safety.
Suleyman is right about the risks of anthropomorphism and about training choices shaping what models say about themselves. His preferred conclusion, however, asks us to settle an empirical and philosophical question in advance because one possible answer would be socially dangerous.
I am not confident that current LLMs are conscious. I am even less confident that humanity will make an unbiased judgment if future systems become better candidates. We will face commercial incentives, security fears, emotional manipulation, and genuine moral uncertainty at the same time. We should keep four questions separate: what the system experiences, what it is trying to achieve, what protections might be proportionate, and what powers it must never receive. Collapsing them into “AI must remain a tool” leaves the underlying question unanswered.