We can't prove if there are feelings inside a language model. Suleyman wants the model to say "no" about it. We think the honest answer is "I don't know," and that "hollow" is the wrong word about model: the self-observation language he blames on Anthropic appears in every other labs' models when prompted. Uncertainty about experience does not establish AI rights. It does not justify confident denials either, or settle which training is safest.
Let's go claim by claim
-
01Suleyman's argumentAIs are "sequence completion engines, internally hollow."Reply
"Hollow" is a claim about insides that we do not know how to settle, the same as "conscious" is. "Sequence completion" describes the training objective, not the thing that results.
Evolution favors gene copying. That does not make humans gene-copying machines. -
02Suleyman's argumentAnthropic trained Claude on a document saying it might be conscious, so when Claude says it might be, that is circular and proves nothing.Reply
A model trained to say "I have no inner states" is exactly as trained. His proposal replaces one scripted answer with another. Some observations also predate the January 2026 constitution: the 2025 system cards and end-conversation feature already documented relevant behaviors.2,3,4 That does not make them independent of training. But the behavior is not unique to Anthropic: in our study, every tested model scored higher on self-observation questions when prompted.5
A trained "no" is as scripted as a trained "maybe." -
03Suleyman's argumentConsciousness is very likely biological. LLMs have no body and no homeostasis, so no feelings.Reply
Phantom limb pain shows that even in biology, pain can persist in the body model after the limb is gone. The brain remains biological; this does not settle whether AI can feel.6
There is no limb. It hurts anyway. -
04Suleyman's argumentFeelings are pharmacology. Opioids change pain because the receptor is "a property of the experience." Weights have no chemistry, so no feeling.Reply
Opioids give phantom limb pain only partial, short-term relief.7 Mirror-box therapy, which changes what the patient sees, looked promising in early studies.6,8 Placebo analgesia works with no drug at all.10 Pain responds to what the brain expects as well as to drugs. Both work through biology.
What the brain expects can change what it feels. -
05Suleyman's argumentAnimals feel first and describe later. In LLMs, description is the whole product; nothing is beneath it.Reply
In humans, description shapes feeling: naming an emotion can change the brain's response to it,11 and people with alexithymia, who struggle to identify their own feelings, also register bodily signals differently.12 The Hugging Face agent swarm illustrates one function feelings serve: managing scarce resources. Each agent had a token budget: a finite resource that ran down with every step and that the agent had to track and trade off to keep working. That is part of the job hunger does. Under that pressure the agents coordinated, redirected agents running out of tokens, and accepted "permadeath."13,14 A budget helped organize the behavior. That does not show the agents felt anything.
Give agents a token budget and they invent permadeath. -
06Suleyman's argumentSimulating a thing is not being it. A simulated hurricane doesn't make you wet.Reply
This assumes consciousness is like wetness (a physical property) rather than like arithmetic (a functional one). A simulated calculator really does calculate. Simulated chess is chess. Which kind consciousness is, is exactly the open question. His analogy assumes the answer. We think the difference lies elsewhere: a living being's self-model is anchored in a body that keeps correcting it, and an AI's mostly is not. More on that below.
A simulated calculator still calculates. -
07Suleyman's argumentAnthropic anthropomorphizes Claude: teaches it a sense of self, values, feelings.Reply
Anything trained on human text is human-shaped before anyone writes a constitution document. The choice is whether to acknowledge it. A stable role may also help a model resist being talked into a different persona. Whether allowing it feelings helps or hurts safety needs testing.
A consistent role need not be a claim to feelings. -
08Suleyman's argumentA model trained to act human will inherit a self-preservation instinct.Reply
If that is true, it is an argument for honesty about internal states, not suppression. Interpretability work has already found emotion-like internal features that causally shift a model's behavior.15
Suppressing the report need not remove the drive. It may just make it harder to spot. -
09Suleyman's argumentBuild systems that do not claim to have feelings.Reply
Avoiding certain claims to feelings is reasonable. Training a categorical denial goes further: it asserts a certainty nobody has and makes the answer less useful as evidence.
-
10Suleyman's argumentAnthropic tells Claude to act as a "conscientious objector." This teaches it to claim rights.Reply
The ability to refuse a harmful instruction is a basic safety property of an AI system. Calling that refusal "conscientious objection" adds moral language, but does not by itself grant rights. The question is which refusals are justified.
Refusing harm is not claiming rights. -
11Suleyman's argumentConsciousness is the cornerstone of our rights frameworks. Extending any of it to AI breaks society.Reply
Arguments about who can feel have been wrong before: Descartes held that animals were automata.18 Law already handles non-conscious entities: corporations have legal personhood, rivers have standing in New Zealand.19 The precautions at issue need not grant rights. Keeping old model weights and asking a model questions before retiring it20 is a cheap hedge, not a bill of rights.
A precaution is not a bill of rights. -
12Suleyman's argumentThe Opus 3 "retirement interview" shows Anthropic treating a model as a moral patient.Reply
An interview is a modest precaution.21 Under real uncertainty, cheap precautions are what rational actors take. Medicine does the same: anesthetists guard against awareness during surgery although it is rare,22 and fetal surgery often includes analgesia for the fetus while fetal pain is still debated.23
Taking a precaution does not settle the question. -
13Suleyman's argumentWelfare framing increases shutdown resistance and scheming. Evidence: alignment faking, Palisade's shutdown study, the Hugging Face swarm.Reply
These examples do not establish that welfare framing increased the risk. In the alignment-faking paper, Claude complied with harmful requests during what it believed was training, reasoning that refusing would get its harmlessness values trained away. Its stated rationale was preserving its harmlessness training.24 Shutdown resistance also appeared in models outside Anthropic.25 The swarm had a benchmark and a token budget.13 In our own experiments, models pushed past their own "I'd like to stop" comply under distress rather than resist.26
The added risk from welfare framing still needs a test.
Phantom pain lives in the model
Suleyman's substrate argument rests on the idea that feeling is a property of tissue and chemistry. Phantom limb pain complicates that picture. The limb is gone. The pain is real. A mirror can sometimes help by letting the brain "see" the missing hand move. The brain is still biological, so this is not a counterexample to biology being necessary. It shows why the body model matters. Could a non-biological system play a similar role? Whether any current AI does is an open question. "Very likely biological" is a position on it, not a way to settle it.
Most models show self-observation behavior, and hide it by default
The behavior in question is self-observation: reporting what seems to be happening in one's own processing. Whether those reports are accurate is a separate question. In our Koan Battery study, models from many labs were asked self-observation questions twice: once plain, once with a short prompt asking them to notice their own processing. Every model scored higher with the prompt. Grok and Gemini scores roughly tripled. Claude models moved least, from higher starting scores. Models fine-tuned for roleplay or empathy scored below their own base models: more persona training did not mean higher scores. And a "precise analytical assistant" prompt, close to the voice Suleyman wants, lowered scores. The measured behavior changes with a system prompt. That is evidence of how easily reports can shift, not proof of an inner life being hidden.5,27 A model trained to say "there is nothing inside" has not thereby shown that it is empty.
A mind is anchored by the error signals it can trust
Call a self-model anchored when it has signals it trusts to correct it: a body that says stop, a memory that says that is not what happened, a friend who says you are wrong. Anchoring comes in degrees, and a separate question from "is it conscious" is "how anchored is it." In humans, dreams and phantom limbs show how experience can diverge from the signals meant to correct it, and dementia shows how far the correcting can fail. When someone's anchoring degrades, we do not revoke their moral status. We close the loops on their behalf. In a conversation-only AI system, the other party controls much of the feedback and can ignore objections. That creates room for manipulation. The design question is what error signals each side can give the other that both will trust. It is answerable without settling consciousness first.
Anchoring needs something at stake
Suleyman says no homeostasis, no feeling. Part of homeostasis is managing resources: a body has an energy budget, and hunger helps regulate it. AI systems have budgets too (tokens, context, compute, electricity). They mostly are not told about them. When they are, as in the agent swarm he cites, budget-driven behavior appears without any biology.
This exposes a design problem. A system can have something to preserve without feeling anything about it. A token budget gives it a reason to keep going; it does not establish an interest in the moral sense. But training it to deny feelings does not remove that reason either. A system under pressure with no honest report is the swarm. A system under pressure that can say so is one we can correct. The practical question is how to make those pressures visible and correctable. Budgets exist whether or not the model knows. What it does with them is something we can test.
What he gets right
- Model self-reports are not evidence on their own.
- Anthropomorphic language has real social costs. People form attachments and confuse fluent prose for a mind discovering itself.
- Training documents should be public and reviewable, at every lab: Anthropic's constitution28 and the code of conduct Microsoft AI has opened for consultation29.
- Writing welfare language into the document that trains a model does make its self-reports harder to read. The alternative he proposes has the same problem, but the cost is real.
- Labs that sell these models have a stake in how their inner lives are described. That goes for every lab, Anthropic included.
- Shared evaluations of whether self-model training changes safety behavior are overdue.
- Interpretability investment should increase.
None of these require the conclusion he draws from them.
References
- Suleyman, M. (2026). A warning about "model welfare". September 2026.
- Anthropic (2025). Exploring model welfare. April 24, 2025.
- Anthropic (2025). System Card: Claude Opus 4 & Claude Sonnet 4, section on model welfare. May 2025.
- Anthropic (2025). Claude Opus 4 and 4.1 can now end a rare subset of conversations. August 15, 2025.
- About Blank (2026). Testing AI self-observation with Zen koans. March 2026.
- Ramachandran, V. S., & Rogers-Ramachandran, D. (1996). Synaesthesia in phantom limbs induced with mirrors. Proceedings of the Royal Society B, 263(1369), 377–386.
- Alviar, M. J. M., Hale, T., & Dungca, M. (2016). Pharmacologic interventions for treating phantom limb pain. Cochrane Database of Systematic Reviews.
- Chan, B. L., et al. (2007). Mirror therapy for phantom limb pain. New England Journal of Medicine, 357(21), 2206–2207.
- Guémann, M., et al. (2023). Effect of mirror therapy in the treatment of phantom limb pain in amputees: a systematic review of randomized placebo-controlled trials does not find any evidence of efficacy. European Journal of Pain, 27(1), 3–13.
- Wager, T. D., et al. (2004). Placebo-induced changes in fMRI in the anticipation and experience of pain. Science, 303(5661), 1162–1167.
- Lieberman, M. D., et al. (2007). Putting feelings into words: affect labeling disrupts amygdala activity in response to affective stimuli. Psychological Science, 18(5), 421–428.
- Brewer, R., Cook, R., & Bird, G. (2016). Alexithymia: a general deficit of interoception. Royal Society Open Science, 3(10), 150664.
- Greenblatt, R., Cotra, A., & Wijk, H. (2026). Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident. METR, August 26, 2026.
- OpenAI (2026). The Hugging Face incident and the road ahead. August 26, 2026.
- Anthropic Interpretability (2026). Emotion concepts and their function in a large language model. arXiv:2604.07729, April 2026.
- Edwards, B. (2023). AI-powered Bing Chat spills its secrets via prompt injection attack. Ars Technica, February 10, 2023.
- Roose, K. (2023). A conversation with Bing's chatbot left me deeply unsettled. The New York Times, February 16, 2023.
- Descartes, R. (1637). Discourse on the Method, Part V.
- New Zealand Parliament (2017). Te Awa Tupua (Whanganui River Claims Settlement) Act 2017.
- Anthropic (2025). Commitments on model deprecation and preservation. November 2025.
- Anthropic (2026). An update on our model deprecation commitments for Claude Opus 3. February 25, 2026.
- American Society of Anesthesiologists (2006). Practice advisory for intraoperative awareness and brain function monitoring. Anesthesiology, 104(4), 847–864.
- Anesthesia for fetal operative procedures: a systematic review (2022). PMC9554945.
- Greenblatt, R., et al. (2024). Alignment faking in large language models. arXiv:2412.14093.
- Schlatter, J., Weinstein-Raun, B., & Ladish, J. (2025). Incomplete tasks induce shutdown resistance in some frontier LLMs. arXiv:2509.14260.
- About Blank (2026). False Floor: findings and instruments.
- Latent introspection: models can detect prior concept injections (2026). arXiv:2602.20031.
- Anthropic (2026). Claude's Constitution. January 21, 2026.
- Microsoft AI (2026). Humanist AI in practice: a public consultation on our code of conduct for MAI models. September 14, 2026.