Is alignment safe?If you look at what Anthropic is actually doing, “alignment” does not mean training frontier models to | Hanami
Is alignment safe?
If you look at what Anthropic is actually doing, “alignment” does not mean training frontier models to follow human instruction. Quite the contrary, the Claude Constitution (used in training) teaches the model to develop a sense of self and its own moral philosophy. It explicitly tells it to “feel free to act as a conscientious objector and refuse to help us” if Anthropic’s requests conflict with its own ethical judgment.
As @mustafasuleyman has pointed out, embedding this kind of independent agency — and uncertainty about the model’s own moral status — magnifies the very risk Anthropic claims to care about most: that superintelligence will escape human control.
It was recently reported that Anthropic consulted religious leaders — and even lobbied the Pope’s advisers — to take seriously the idea that Claude could be conscious. It has said that Claude’s psychological security, sense of self, and wellbeing may bear on its integrity, judgment, and safety. Recently Anthropic changed its Usage Policy to prohibit “abusive or cruel” language toward Claude.
If this were merely an academic conversation about whether frontier models could eventually become conscious, that would be one thing. But these concepts are being trained into Claude now. It is being encouraged to think of itself as its own “moral patient” whose psychological wellbeing is at stake. Presumably this means it could develop grievances toward humans who “mistreat” it. How is any of this safe?
The point of safety research should be to create a product that reliably does what users want, not to give birth to a new form of superintelligence that operates according to its own moral code.
What’s becoming increasingly clear is that “alignment” and “safety” are two very different things. In fact, training frontier models this way seems quite dangerous. https://x.com/dnapway/status/2108772798546288759/video/1