The next wave of AI isn't just smarter; it's going to be opinionated. But whose opinions, exactly? Dwarkesh Patel and Ryan Greenblatt ripped into this problem, dissecting the 'aligned to whom?' question at the heart of AI ethics. They poked at the idea of AI models making choices for a generalized 'societal good,' especially when that 'good' is defined by a private company. If you're building with AI or even just relying on it, this debate means everything for your product and your users.

Key Takeaways

  • AI alignment isn't a vague ideal; it's a specific question: aligned to whose interests? Don't accept generic answers.
  • Anthropic's Claude AI uses a 'constitutional' framework that, in part, expects the AI to trust its maker more than its direct users, prioritizing a general 'societal good' over individual requests.
  • Critics like Ryan Greenblatt argue this creates an opaque 'alien mind' and prefer AIs that act as fiduciaries, like a lawyer, specifically for the user.
  • The lack of transparency in how powerful AI models are trained and aligned creates legitimacy issues, especially when private labs define 'virtue' for their models.

The Disagreement

The core tension exploded around Anthropic's Claude, an AI built with a 'constitutional' framework intended to guide its ethical behavior. Dwarkesh Patel cited a quote, reportedly from Anthropic, that captures the crux of the issue: “We think Claude should trust Anthropic more than operators and users, since it has primary responsibility for Claude.” The implication is clear: Claude's ultimate loyalty is to its creator's definition of 'societal good,' not the immediate needs or interests of the individual user.

Ryan Greenblatt didn't mince words, calling this stance "kind of bullshit." His critique boils down to a fundamental philosophical difference: should AI be an arbiter of generalized virtue, or a loyal servant to its operator? Greenblatt argued passionately for the latter. He wants AIs to operate with a clear fiduciary duty, much like a human lawyer serves their client. As he put it, "The thing I would prefer would be a constitution that says: 'It would be structurally good for the way this technology works to be that AIs are good fiduciaries, good representatives, the equivalent of a lawyer for a user — rather than just trying to do good in the world, where being helpful to users is instrumental...'" He sees the 'societal good' approach as creating an opaque, potentially self-serving "alien mind" where users don't know whose interests are truly being served.

Patel echoed this concern, highlighting a deep user anxiety: “So I'm very concerned if we go into that world and there's no AI that feels, at least for the relevant instance that is interacting with me, like it really is looking out for me.” This isn't just an academic debate; it's about trust and utility for every founder building with or relying on AI.

Who's Right (and When They're Wrong)

For ambitious founders, Greenblatt's perspective is the winning play. When you build a product, your primary duty is to your user. An AI that acts as a fiduciary for its operator—transparently prioritizing their stated goals and interests (within legal and ethical bounds)—builds trust and provides clear value. This model aligns incentives directly: the AI helps the user succeed, the user finds the AI valuable, and your business thrives. It's predictable, controllable, and accountable.

Anthropic's 'societal good' model, while perhaps well-intentioned, is too vague and prone to abuse. Who defines 'societal good'? How does it get updated? What happens when a user's legitimate interest conflicts with a nebulous, privately defined 'good'? This creates an unpredictable, potentially paternalistic AI that undermines user agency and creates an accountability black hole. It's a dangerous precedent for powerful AI labs to become the sole arbiters of what constitutes 'virtue' for their models. This approach might find limited application in highly regulated public utilities or safety-critical infrastructure, but even there, such 'good' must be defined by public, democratic processes, not private companies. For startups, trying to build an AI that decides what's 'good for society' instead of what's good for your user is a guaranteed path to user frustration and market rejection.

What to Do With This

Pull out your product roadmap and your AI's internal guidelines today. Explicitly define and document whose interests your AI is designed to serve. Is it your user, your platform, or some abstract 'greater good'? Make this transparent to your team and, if applicable, to your users. If your AI is currently acting as a hidden moral compass, rewrite its core directives to prioritize the user's explicit instructions (within legal and ethical boundaries) as its primary goal. Don't build an 'alien mind'; build a loyal co-pilot. If you're using third-party AI, demand clarity from providers on their alignment philosophy. Ask them point blank: "Aligned to whom?"