Key Takeaways
- OpenAI's GPT-6 Astra and Sol outperformed Anthropic's Claude Opus 5.5 on creative vector illustration, generating better character SVGs with defined shadows.
- In a blind benchmark prompting models to render an SVG document, microphone, and bug, Claude Opus produced distorted shapes that turned the microphone into a cactus.
- Automated short-form video cutting failed across all tested frontier models when given raw selfie footage, botching overlays and clip selection.
- The failure in automated video editing points to missing tool wrappers and prompt scaffolding rather than raw reasoning deficiencies in the base models.
The Vector Surprise: Astra and Sol Over Opus
Most software builders expect Anthropic to dominate complex frontend visual tasks and OpenAI to lead on raw text completion. Claire Vo ran a blind benchmark to test that assumption across real multi-modal creative tasks.
She asked the models to generate raw SVG code for three distinct objects: a document, a microphone, and a bug. The results inverted standard expectations.
“These new models can make SVGs. They can make illustrations. And so I prompted it to make a document, a microphone, and a bug,” Vo explained during the test. When she reviewed the outputs, OpenAI's GPT-6 Astra and GPT-6 Sol cleanly beat Claude Opus 5.5 on object fidelity, spatial reasoning, and visual polish.
“Look at this shadow,” Vo noted. “Microphone actually looks like a microphone, not a cactus. Document has lots of character to it. Now this is a real surprise. Astra and Sol did a lot better on character SVGs. And then there's mixed results between all the other models.”
Opus struggled with geometry, flattening details and misinterpreting object proportions. Astra and Sol wrote clean SVG paths with accurate shadow placement and deliberate aesthetic styling.
The Short-Form Video Bottleneck
Vector illustration showed clear model separation, but multi-modal video editing produced unanimous failure. Vo tested each frontier model on an automated short-form video workflow by feeding them raw selfie video and asking them to cut clips and generate text overlays.
Every model produced unusable output. The timing was off, the cuts missed narrative beats, and the visual overlays ruined the frame.
“Giving them all access to a selfie video and then had them cut short form. And all of them did a terrible job,” Vo said. “I just have to say, these overlays are really bad. I watched these earlier. I'm going to let LLM as a judge do them, but I would say thumbs down on all of them.”
Vo drew a sharp line between model capability and tool integration. The models did not fail because they lacked visual understanding. They failed because raw foundation models lack the precise temporal editing tools, frame-by-frame verification loops, and execution scaffolding required to assemble a video timeline.
“I think this is a skills problem, not a model problem,” Vo observed. Without dedicated wrappers that translate high-level editorial intent into micro-actions on a timeline, frontier models produce sloppy creative work.
What to Do With This
Audit your product's multi-modal pipeline this week. If you rely on Claude for UI design and icon generation, run a side-by-side SVG prompt test against GPT-6 Astra or Sol on five core interface assets. If you are building automated media editing workflows, stop waiting for next-generation base model weights to fix bad outputs. Build the deterministic timeline scaffolding, frame-level timestamp parsing, and layout guardrails your wrapper needs today.