Claire Vo's 'Barbie Bench' Exposes AI Spatial Blindspots
Claire Vo explains why Claude Opus 5.5 and GPT-6 fail the Barbie Bench 3D test, proving frontier AI still cannot render basic human anatomy.
10+ hours of podcasts, in 5 minutes.
Claire Vo runs a live, blind benchmarking session comparing OpenAI's GPT-6 Sol and GPT-6 Astra against Anthropic's Claude Opus 5.5 across real-world workflows including PRDs, email triage, frontend coding, SVGs, and video editing. Vo evaluates the models on speed, ergonomics, personality, and design aesthetic, contrasting her human ratings with automated LLM-as-a-judge scores. The episode highlights why Claude Opus 5.5 excels at complex B2B frontend and agentic tasks while OpenAI's models dominate in writing clarity, token efficiency, and creative illustration.
Claire Vo explains why Claude Opus 5.5 and GPT-6 fail the Barbie Bench 3D test, proving frontier AI still cannot render basic human anatomy.
Claire Vo shares her blind taste test framework for testing frontier AI models like Claude Opus 5.5 and GPT-6 Astra across real workflows.
Claire Vo benchmarks Claude Opus 5.5 against GPT-6 Sol across speed, price, and verbosity in daily product workflows.
Claire Vo blind-tests GPT-6 Astra, Sol, and Claude Opus 5.5 on SVGs and video editing. Here is why the results surprised everyone.
Claire Vo benchmarks Claude Opus 5.5 against GPT-6 Sol and Astra on frontend code, exposing key design biases and UI bottlenecks.
Claire Vo found automated LLM judges reward verbose outputs over human readability. Here is how human blind tests expose model evaluation flaws.
10+ hours of podcasts, distilled into one 5-minute read. Free, every Sunday morning.
One email a week. Unsubscribe with one click.