Key Takeaways

  • Claude Opus 5.5 fixes the preachy tone and verbosity of prior releases, earning its way back into daily engineering stacks.
  • Claire Vo assigns Opus 5.5 to frontend UI prototyping, SVG generation, architecture reviews, and adversarial pull-request checks.
  • OpenAI Codex remains Vo's default choice for raw computer use, desktop workflows, and tool execution.
  • Autonomous video production remains a weak spot for Opus 5.5; when paired with ElevenLabs connectors and MCP to cut TikTok shorts, it produced low-quality edits.
  • Opus 5.5 suffers from noticeable latency compared to faster alternatives like Soul and Astra.

Pick Models by Cognitive Task, Not Loyalty

Most engineers still try to force a single AI model to run their entire development cycle. Claire Vo abandoned that approach after testing Anthropic's Claude Opus 5.5 alongside OpenAI Codex. Vo had previously stepped away from Claude because earlier models delivered preachy, overly verbose responses. Opus 5.5 corrected those tone issues and tightened its safety boundaries, but it did not replace Codex across the board.

Instead, Vo splits her stack based on task type. Models that excel at high-level reasoning and visual generation stumble when given mechanical control over the desktop. Trying to make one tool do both wastes time and degrades output quality.

The Claude Sweet Spot: Architecture and Critique

Opus 5.5 shines when placed in an advisory or creative design seat. Vo relies on the model for three core areas: frontend prototyping, architectural design decisions, and adversarial pull-request reviews.

“Where am I using it a lot practically? I'm using it in PR review. I'm using it in architecture questions,” Vo explained. “And then I'll be using it in front end.”

In frontend development, Opus 5.5 excels at creating working interfaces and generating clean SVG assets from scratch. In code review, its strength lies in spotting edge cases and challenging architectural assumptions before code merges. The model acts as a rigorous technical peer rather than a silent auto-complete tool. However, speed remains a friction point. “That being said, it felt consistently slower still than Soul or even Astra,” Vo noted regarding the model's response times.

Where Computer Use and Media Automation Fail

When tasks move from code review to direct execution, the stack shifts back to Codex. Vo keeps Codex as her primary engine for general desktop workflows and computer use.

“Where am I still not going to the claw models computer use? I just think Codex is so much better cutting videos,” Vo said.

That difference became obvious during automated media experiments. Vo attempted to build an automated editing pipeline using Model Context Protocol (MCP) integrations alongside an ElevenLabs connector. The goal was simple: take raw selfie videos and automatically edit them into TikTok-style vertical shorts.

Opus 5.5 failed to handle the assignment effectively. “I've been using the ElevenLabs connector and MCP to cut selfie videos into TikTok style shorts,” Vo observed. “It just had both terrible taste.”

The failure highlights a real limitation in current model intelligence. While Opus 5.5 can reason through system designs on paper, it lacks the pacing and aesthetic judgment required for media editing tasks. For low-level operating system tasks and execution loops, Codex remains the superior platform.

What to Do With This

Audit your engineering workflow by separating creative reasoning from automated execution. Route your architecture RFCs, SVG asset generation, and PR reviews to Claude Opus 5.5 to catch structural bugs early. Route script execution, terminal commands, and desktop agent workflows to OpenAI Codex.