Key Takeaways

  • Public leaderboards measure synthetic trivia, while production work depends on ergonomics, speed, and day-to-day readability.
  • In head-to-head testing between Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Astra and Sol, Opus 5.5 led on complex frontend UI and agent tasks, while OpenAI models won on concise writing and token efficiency.
  • Automated LLM judges often diverge from human preference because programmatic rubrics miss formatting annoyances and bloated phrasing.
  • Testing frontier models requires blinded evaluations scored from the viewpoint of the person receiving the work rather than the prompter.
  • Engineering teams can deploy Claire Vo's Blind Taste Test Benchmarking Method to choose daily driver models based on actual workflow utility.

Claire Vo's Blind Taste Test Benchmarking Method

Claire Vo rebuilt her evaluation bench when fresh models hit the market. As Vo explained, “I did think it was an important moment. Right now, we've had Astra come out, Fable's been refreshed. I just wanted to completely rethink the How I AI bench.” Instead of comparing mismatched tiers, she runs structured, blinded tests across five steps:

  • Step 1: Group by Capability Tier: Group models into like-for-like capability tiers (e.g., matching flagship models against flagships rather than pairing lightweight models with frontier systems) to ensure fair comparison.
  • Step 2: Assign Realistic, Multi-Discipline Prompts: Run both directed (tight constraints) and open-ended (one-shot) real-world tasks across PRDs, inbox triage, frontend UI, backend spec auditing, long-running agent research, and creative outputs.
  • Step 3: Score Blindly on Practical Ergonomics: Evaluate anonymized outputs on a 1-to-5 scale from the perspective of an end-user recipient, penalizing structural annoyances (e.g., overused em dashes, dense unformatted text) and rewarding legible, actionable outputs.
  • Step 4: Contrast Human Vibe Ratings with LLM Judge: Run automated evaluations through an LLM judge alongside human scores to identify divergence between formal programmatic metrics and practical human preference.
  • Step 5: Unmask and Synthesize Model Patterns: Feed anonymized scores and notes back into an analysis prompt to identify overarching strengths, failure modes, cost-efficiency trade-offs, and personality quirks across model families.

Vo keeps the grading fast and practical: “I give a vibe one through five and I give terrible notes like simple and straightforward. No complaints. I give it a task. I look at it. I evaluate it as if I am the recipient of this task, which is common. And then I rank it.” Once the run completes, she uses Claude to aggregate her raw notes: “And then what I do is I give it to Claude and I say, 'Claude, what do I think?' Analyze. And we're going to see if my prediction stands right.”

When This Works (and When It Doesn't)

This method works when selecting daily driver models for production teams and knowledge workers where synthetic benchmark scores do not reflect actual ergonomics, speed, and real-world usefulness. It exposes practical friction points, like bad formatting or unreadable code, that automated evals overlook.

It breaks down when you need strict mathematical determinism, low-latency API benchmarking, or regulatory compliance auditing. Blind vibe checks cannot replace unit tests for deterministic data pipelines or security audits for code generation.

What to Do With This

Pick the top three tasks your team runs through AI every week, such as messy meeting notes turned into PRDs, customer ticket triage, and React component prototyping. Set up a spreadsheet with columns for Model A and Model B, stripping model names from the raw outputs.

Have a teammate paste outputs from Claude Opus 5.5 and GPT-6 Astra side by side without labels. Score each output from 1 to 5 based strictly on how much editing you must do before sending it to a colleague. Feed the raw scores into an LLM at the end of the week to reveal which model actually saves your team time.