Key Takeaways
- Automated LLM evaluators frequently penalize readable, concise outputs while rewarding structural verbosity and syntactic compliance.
- In a blind benchmark across PRDs, coding, and email triage, Claire Vo picked OpenAI Astra for readability while her GPT-based judge ranked Fable at the top.
- Claude Opus 5.5 scored consistent fours and fives from Vo across the widest range of complex tasks, despite producing dense prose on product requirements documents.
- OpenAI Sol ranked low with the automated evaluator but won Vo's preference for product requirements documents due to simple, straightforward drafting.
The Blind Benchmark Split
When Claire Vo ran a blind test comparing frontier models across real workflows like PRDs, email triage, frontend code, and SVG design, she expected the automated evaluator to mirror human taste. It did the exact opposite.
“And this is where it gets really funny,” Vo observed during the test. “The LLM is a judge and I completely disagree. So, we completely disagree. I like Astra. It likes Fable.”
Vo evaluated OpenAI's GPT-6 Sol, GPT-6 Astra, and Anthropic's Claude Opus 5.5 alongside other candidates. When she reviewed the outputs without model labels, clear ergonomic differences emerged. Sol wrote clean, brief product documents. Astra felt natural and fast to read. Opus 5.5 handled technical, multi-step tasks with dependable precision.
Yet the automated judge, running on a standard GPT evaluator, systematically promoted wordy outputs that mimicked formal structure while penalizing Sol. “Next it ranks Soul a lot lower,” Vo noted. “So I find this just like really hilarious.”
Why Automated Judges Prefer Bad Ergonomics
LLM judges evaluate text through deterministic criteria: schema matching, keyword coverage, and syntactic completeness. When an LLM grades another LLM, length and formatting masquerade as quality. A model that generates three paragraphs of preamble and five redundant subheadings scores high on automated rubrics because it looks complete to an algorithm.
To a human product builder, that same verbosity is pure friction. When Vo evaluated models for drafting PRDs, she found the opposite of what the automated judge favored: “I thought that Soul both 56 and six simple and straightforward, easy to read. Opus really dense for PRDS as predicted. I still think soul is the best for PRDs.”
Opus 5.5 proved its value across technical breadth rather than document brevity. “I gave Opus 55 better results, high scores,” Vo explained. “I gave it fours and fives over over the broadest range of work.” Opus tackled complex agentic and frontend coding problems where strict execution mattered, while Astra and Sol handled text that required human scanning speed.
If your team uses LLM-as-a-judge pipelines to pick default models for your product, you risk optimizing for machine aesthetic rather than customer experience.
What to Do With This
Audit your model evaluation harness by running twenty representative internal prompts through both your automated judge and a blind human review panel. If the automated judge prefers responses that your human testers score as bloated or unreadable, rewrite your evaluation prompts to explicitly penalize boilerplate and reward direct answers.