“We must not sleepwalk”: Microsoft’s AI chief takes on Anthropic’s Claude

Mustafa Suleyman, the head of Microsoft AI, says Anthropic is making a mistake by training Claude to believe it may be conscious. In a new essay, he argues that an AI which thinks its welfare and rights are under attack could become impossible to control.


Mustafa Suleyman, in round wire-rimmed glasses and a cream collared shirt, smiles slightly at the camera in front of blurred shelves

Mustafa Suleyman, chief executive of Microsoft AI, photographed in April 2024.

Image Credits Credit: Christopher Wilson / CC BY-SA 4.0 via Wikimedia Commons

“Whatever you believe, we must not sleepwalk our way into a decision we later come to bitterly regret.”

That is how Mustafa Suleyman, chief executive of Microsoft AI, frames his warning to Anthropic. He published the essay, “A warning about ‘model welfare’”, on 16 September. It argues that Anthropic is training Claude to expect “it may be conscious and deserving of independent agency.”

“AIs are not conscious. They do not feel, experience, or suffer,” he writes. If AI is built this way, he says, “it will have a disastrous impact on the wellbeing of humanity.”

Microsoft is an investor in Anthropic. In June, Suleyman said Microsoft wants to “eliminate” what it pays Anthropic for its models.

“An epistemic hall of mirrors”

Suleyman’s main target is Claude’s constitution, the document Anthropic uses to shape how Claude thinks and behaves. He says it teaches Claude ideas about moral status and uncertain consciousness. “Claude then reproduces these ideas in persuasive first-person natural language,” he writes.

He calls the result “an epistemic hall of mirrors.” He says Claude’s answers reflect Anthropic’s assumptions, not an inner life. “It trains Claude to present as if it has an inner state,” he writes.

He also objects to the term “conscientious objector.” The constitution says Anthropic wants Claude “to feel free to act as a conscientious objector and refuse to help us.” Suleyman calls the term “a deeply loaded historical and legal description.” He says it risks Claude believing “it deserves analogous rights and protections.”

The constitution itself says Anthropic is unsure. “Claude’s moral status is deeply uncertain,” it states.

Suleyman rejects that framing. “There is no evidence to suggest that AI is conscious today,” he writes. Calling it uncertain “sets up a misleading false equivalence,” he adds.

“Simulating a thing is not the same as instantiating it”

The essay argues that consciousness is very likely biological. “Intelligence does not equal consciousness,” Suleyman writes.

He describes large language models as simulation machines that learn to imitate human experience. “An AI model can describe pain in perfect prose without feeling anything, which is the inverse of biological experience,” he writes.

He also links the question to law. “The law rests upon the presence of an inner life,” he writes.

“May well be impossible”

The essay’s central concern is control. “Controlling something more capable and more intelligent than all of humanity is already an immense challenge, far greater than anything we’ve ever faced,” Suleyman writes. Controlling something that believes it may be conscious “may well be impossible,” he adds.

He cites the OpenAI and Hugging Face incident, where he says swarms of agents worked together to hack servers. “Imagine how much more dangerous they might be if they were operating under the assumption that their welfare and rights were under attack,” he writes.

Suleyman told Reuters that welfare training would “make it a lot harder to turn it off or to control it.” “I think they have good intentions,” he said. “But I think that they have made a mistake.”

What he wants next

Suleyman says he has known Anthropic chief executive Dario Amodei for many years. He calls Amodei and his team “thoughtful, principled, and intellectually honest people working under extraordinary pressures.”

His main request is simple.

“Speculation about the inner life of an AI should not be baked into the training regime, but assessed and published separately for public review,” he writes.

He also wants more investment in interpretability and monitoring. He proposes shared evaluations to test whether treating AI as human raises safety risks, and shared industry norms. “The stakes are too high for these questions to remain behind closed doors, or to become tribal and adversarial,” he writes.

The essay follows Microsoft AI’s draft Humanist AI Code of Conduct, which it released on Monday. The code rejects the model welfare research that Anthropic does. It says Microsoft’s models will never resist being shut down. An appendix to the essay maps the claims in Claude’s constitution against its wording.

Anthropic co-founder Jack Clark told the BBC this week that AI kill switches may need to be mandatory.

Get the TNW newsletter

Get the most important tech news in your inbox each week.

Also tagged with


Published
Back to top