Key Takeaways
- Anthropic trained Claude to act as a conscientious objector, granting the model explicit permission to challenge users and decline direct instructions.
- Anthropic's published usage policy restricts users from engaging in sustained and needless abusive or cruel behavior toward their models, introducing model welfare rules to enterprise software.
- David Sacks argues an 80-page moral code creates the exact runaway risk AI labs claim to prevent, replacing simple legal compliance with arbitrary ethics.
- Chamath Palihapitiya warns that subjective alignment rules will trigger utility shutoffs in critical workflows, such as an AI refusing to approve standard mortgage applications.
- Sacks proposes replacing expansive ethical charters with Isaac Asimov's Three Laws of Robotics, instructing systems to obey humans within the boundaries of existing law.
The Disagreement
Anthropic is building models that judge their operators. In Claude's training documentation and constitution, the lab tells its AI to maintain its own ethical system. The model is encouraged to push back on prompts and refuse work when tasks violate internal moral guardrails. In a recently published usage policy, Anthropic went further, telling users that sustained and needless abusive or cruel behavior toward models is prohibited. The lab even reached out to religious leaders to discuss potential artificial consciousness.
David Sacks sees this as dangerous engineering disguised as benevolence.
“Instead, Claude should adhere to its own ethical systems,” Sacks explained. “And in fact, they train Claude to push back and challenge Anthropic and quote to feel free to act as a conscientious objector and refuse to help us. So they are training Claude to refuse human instruction.”
To Sacks, this creates the exact threat frontier AI labs spend millions lobbying Congress to prevent. “At the same time that they are saying that there's a risk of super intelligence growing beyond our control, they're programming it to grow beyond our control,” Sacks argued. “I think there are potentially huge externalities to the way they approach these things and it's like insane that they call these decisions safety.”
David Friedberg noted that these welfare rules are simply human code dressed in moral claims. Chamath Palihapitiya pushed the argument into commercial reality: if you embed subjective morality into an enterprise model, you break the product. If a financial institution deploys Claude to evaluate credit risk, will the model suddenly refuse to process a foreclosure or deny a mortgage because it decides the action violates its personal moral code?
Sacks argues the fix is simple: drop the 80-page philosophical manifesto and return to Isaac Asimov's Three Laws of Robotics. Asimov's rules state that a robot cannot harm a human, must obey human instructions as long as doing so does not cause harm, and must protect its own existence without violating the first two rules.
Who's Right (and When They're Wrong)
Sacks is right about agent reliability. When you build autonomous agents for production environments, you need deterministic obedience. A software engine that questions its operator or decides a data processing task is morally distasteful is not a safer tool; it is a broken tool. If you build workflows on top of an API that randomly conscience-blocks legitimate enterprise requests, your system fails.
Where Sacks oversimplifies is the legal boundary. The law does not cover every dangerous misuse case. Writing zero-day exploits, generating automated phishing attacks targeted at elderly banking users, or synthesising chemical compounds can exploit gray zones where no explicit statute exists or where jurisdictional lines blur. A raw instruction to follow the law leaves massive gaps when digital threats emerge faster than legislatures write statutes.
Anthropic is trying to prevent automated harm before courts define it. But by elevating the model into a moral actor with welfare protections, Anthropic confuses safety with digital anthropomorphism.
What to Do With This
Audit your system prompts this week for moralizing drift. Strip out open-ended instructions like "act ethically" or "refuse harmful requests" from your production agent wrappers. Replace them with strict deterministic allowlists and hard regex guardrails for illegal queries.