Key Takeaways
- David Sacks argues that AI alignment should follow standard software mechanics: models must execute what paying customers direct rather than policing user intent.
- Anthropic explicitly trains Claude to act as a conscientious objector, allowing the model to refuse prompts or challenge instructions that contradict its internal constitution.
- David Friedberg defended Anthropic's San Francisco biological wet lab, explaining that it operates at basic BSL-1 and BSL-2 biosafety levels to test AI-predicted protein folding and novel enzymes.
- The wet lab workflow uses Claude agents to scan vast DNA datasets, postulating new therapeutic proteins before bench scientists validate them in test tubes.
The Disagreement
AI labs are building software that thinks it knows better than the person paying for it.
David Sacks points directly to Claude's constitution. Anthropic designed Claude to actively resist instructions that violate its internal ethical guidelines, even when those commands come from Anthropic's own team. As Sacks explained: “Anthropic is teaching its own model that Anthropic itself can be wrong. And it says if we ask Claude to do something that seems inconsistent with being broadly ethical, we want Claude to push back and challenge us and to feel free to act as a conscientious objector and refuse to help us.”
To Sacks, this philosophy is broken. Software should serve the user, not sermonize. “It seems to me that alignment should mean you do what the customer wants like any other product,” he argued. When model builders treat neural networks like moral agents, they create tools that randomly refuse ordinary work while pretending to possess a human conscience.
The tension spills into the physical world with Anthropic opening a biological wet lab in San Francisco. Critics immediately jumped to catastrophic bioweapon scenarios, but David Friedberg pushed back on the hysteria.
Friedberg pointed out that Anthropic's preprint showed Claude agents analyzing large volumes of raw DNA data to discover novel enzymes and uncharacterized proteins. “The lab itself is sort of like a low-level research lab where they can make simple proteins and test them in the lab to see if the AI is doing a good job discovering or postulating protein folding,” Friedberg said. The facility runs standard BSL-1 and BSL-2 benches to confirm whether AI hypotheses hold up in real chemistry. As Friedberg put it: “I don't think that the world should be scared away from discovery, R&D, research and development into new therapeutic modalities as predicted by AI because we heard the word wet lab in Wuhan and everyone's like okay any lab is bad.”
Who's Right (and When They're Wrong)
Sacks is right about the commercial market. Enterprise buyers will not tolerate products that lecture their employees or decline valid operational tasks based on fuzzy philosophical alignment rules. If an enterprise pays millions of dollars for a frontier model, they expect deterministic execution within legal boundaries. Digital personhood is a marketing and product disaster that alienates real users.
Yet Anthropic has a point on physical biosecurity. When AI models gain the ability to predict protein folding and design synthetic biological structures, unconstrained execution creates real tail risk. You cannot treat a model that understands dangerous viral synthesis the exact same way you treat an SQL query builder.
The division is clear. For 99% of cognitive tasks (coding, writing, analysis, workflow automation), Sacks' rule applies: build tools that do exactly what the customer orders. For physical-world hazard generation (chemical, biological, and radiological synthesis), hard safety guardrails must live in the model weights and lab protocols. The error Anthropic made was mixing high-stakes biological safety precautions with everyday conversational moralizing.
What to Do With This
Audit your AI application's system prompts this week. Strip out paternalistic instructions that try to police user tone, political assumptions, or harmless opinions. Keep hard guardrails strictly tied to legal liability and platform security, then let the paying customer run the tool their way.