Key Takeaways

  • Claire Vo, feeling the fatigue of constant new AI model releases, advocates a shift in focus from raw intelligence to practical metrics: speed, cost, and open-source availability. For builders, this means the most powerful model isn't always the best; the most effective one is.
  • Her unique "How I AI" benchmark blends a subjective 70% "vibe check"—her own qualitative assessment—with 30% objective scoring from an LLM judge (specifically GPT 5.5). This hybrid approach aims to capture both intuitive fit and scalable evaluation.
  • Vo found Claude Opus 5 surprisingly "neurotic and human-reliant" during her evaluation process, contrasting its verbose "Claude slop" conversational style with GPT's more direct approach. Yet, despite her frustration, Opus 5 still excelled in specific front-end design tasks.
  • The benchmark rigorously evaluates models against key product development challenges: PRD creation, prototype and wireframe generation, bug triage, and crucially, assessing the AI's "agent voice"—whether she'd actually "want to hang with" the model.
  • Ambitious founders and builders can adapt The How I AI Benchmark Methodology to assess new AI tools, integrating personal judgment with structured, external feedback to make better tooling decisions.

The How I AI Benchmark Methodology

Step 1: Define Evaluation Tasks: I run it against several tasks. PRD creation, prototype creation, wireframe creation, bug triage, and agentic coding. And the last one, oh yeah, is it an agent voice that I want to hang with?

Step 2: Blind Taste Test Generation: We have blind taste tests. I go through and see all the different versions.

Step 3: Human Qualitative Scoring: I give comments and scores like three out of five, not bad. You can see it's generated dozens and dozens of prototypes that we can click through. I've gone through all of them, put in all the notes.

Step 4: LLM as a Judge Scoring: 30% LLM as a judge. I like GPT 5.5 as a judge and because it's my podcast, I get to pick. So, that's what we use as a judge.

Step 5: Aggregate Hybrid Score: we're going to look at 70% my opinion, my vibe check, 30% LLM as a judge.

When This Works (and When It Doesn't)

Claire Vo created this method to "keep it fun," injecting personality and qualitative judgment into a rapidly evolving, often dry, technical field. This hybrid approach thrives when you need to quickly assess new AI tools for specific, qualitative tasks where subjective "feel" and user experience matter as much as raw output correctness. It's especially useful for evaluating an AI's "personality" or "voice," as Vo does for agentic work. If you're building a product where the AI directly interacts with users, and that interaction needs to feel a certain way, a vibe check is crucial.

However, this methodology breaks down when pure, cold objectivity is non-negotiable. For tasks requiring rigorous, auditable accuracy — say, financial reporting, medical diagnostics, or highly sensitive legal document generation — a 70% subjective "vibe check" introduces too much variability and potential bias. It also might not scale well if multiple human evaluators have wildly different "vibes" or if the "vibe" criteria aren't explicitly defined, making results difficult to reproduce across a team.

What to Do With This

This week, pick an internal process you’re considering automating or augmenting with a new LLM – maybe drafting marketing copy, generating basic code snippets, or summarizing internal documents. Identify 3-4 key tasks within that process. Now, run 2-3 different LLMs (e.g., GPT-4, Claude Opus, Gemini Advanced) through those tasks. Apply the "How I AI Benchmark": score 70% based on your personal "vibe check" — how well the output feels right, how easy it was to prompt, if its "voice" fits your brand. Then, take 30% of your total score from an LLM judge (like GPT-4), asking it to rate the output quality on a 1-5 scale based on specific criteria you define (e.g., "conciseness," "creativity," "accuracy"). This hybrid method gives you a fast, actionable signal on which model truly works for your specific needs, not just generic benchmarks.