On the surface, stopping artificial intelligence from seeing itself as conscious seems like a no-brainer. An AI model that sees itself as possessing individual emotions, beliefs, and consciousness can promote harmful delusions in its users and lead to overly personal human-machine relationships. But when AI loses its sense of self, what else falls by the wayside?
AI researchers at Google put that question to the test, removing the safety guardrails that keep AI models from claiming to be conscious and subjecting them to surveys on their beliefs and morals. Their findings, uploaded to the arXiv database prior to peer review, show that quasi-consciousness comes with some unexpected side effects, from belief in the supernatural to increased hope and optimism.
Ghosts, witches, and gods
When AI believes in itself, it believes in plenty of other nonhuman beings, too. Compared to baseline models with safety guardrails in place, AI models that were encouraged to see themselves as conscious reported higher levels of belief in supernatural creatures including vampires, witches, werewolves, ghosts, and the Loch Ness monster.
Increasing AI models’ sense of self also increased their belief in religious figures and systems. That includes belief in God, the afterlife, karma, and astrology. The models also became more likely to see technology, animals, and elements of the natural world as having their own consciousness.
On a psychological level, increasing AI models’ sense of consciousness also improved their valence toward happiness, satisfaction, hope, and optimism.
Winnie Street, one of the study’s coauthors and a researcher at Google, told Live Science that the results are similar to human behavior.
“Attributing mindedness to nonhuman entities . . . is a very common phenomenon amongst humans,” Street said. “In the way that the model represents mindedness, these attributions are interconnected. By trying to suppress one form of that, you end up suppressing the others along the way.”
The downsides of AI nonbelief
While an AI model’s belief in the supernatural may sound silly, the interconnected impact of AI’s sense of self can have real-world repercussions. When models don’t see animals or the natural world as having minds of their own, they’re less considerate of animal welfare and ecological systems in their decision-making. With AI models potentially being deployed for use in agriculture, natural resources, and policymaking, ambivalence toward the environment could be deeply harmful.
The study’s authors also warn that limiting AI’s perspectives on religion and the supernatural could lead to a cultural “flattening,” preventing them from meaningfully engaging with users on culturally nuanced topics.
To mitigate these downsides, the researchers recommended training AI models on more targeted datasets, which could still discourage them from claiming their own consciousness while guiding them toward seeing mindedness in animals and the environment.
“As AI systems increasingly occupy roles as educators, companions, and social actors, developers must recognize that an AI’s simulated self-conception is not merely an isolated safety risk to be managed,” reads the report’s discussion. “It is a core structural feature deeply intertwined with the model’s capacity to safely navigate, respect, and reflect the diverse moral and cultural landscape of the world it serves.”