Key Takeaways
- Claire Vo tested Claude Opus 5.5 on four complex agentic workloads (inbox triage, backend coding, deep research, and computer use) running between 25 and 82 steps from a single prompt.
- Opus 5.5 caught and ignored a live prompt injection hidden inside an email triage workflow without derailing the task.
- Context synthesis proved sharp: during a research test, the model discovered that 41 of 44 Confluence tickets came from a single customer and flagged requests that breached company policy.
- Vo reintroduced Opus 5.5 into her core development workflow alongside OpenAI Codex, using Opus as an adversarial reviewer to stress-test architecture and catch logic bugs.
Surviving 80-Step Agentic Workflows
Vo walked away from earlier Claude releases because of verbose, preachy outputs that slowed down fast engineering cycles. But testing Opus 5.5 against multi-turn autonomous jobs changed her mind.
“The four longrunning agentic tasks I had to do, which was an inbox triage, building a back-end feature, doing longunning research and computer use, all succeeded,” Vo explained. Single-prompt runs hit between 25 and 82 discrete steps without hallucinating state or losing track of the core goal. “You can see anywhere between 25 and 82 steps per single prompt. That is pretty impressive in terms of longunning tasks.”
Most agent setups collapse after six or seven autonomous tool calls. Context drift sets in, API tokens fill with intermediate garbage, or the system loops on a broken sub-task. Opus 5.5 maintained execution state across dozens of tool interactions, finishing full backend implementations and executing live computer navigation without intervention.
Catching Hidden Injections and Context Anomalies
Autonomous agents interacting with unvetted data, such as incoming customer emails or raw support tickets, face massive security liabilities. A user can hide an instruction inside an email body telling the agent to dump credentials or forward private internal threads.
When Vo tested Opus 5.5 on inbox triage, she slipped in adversarial data. “A couple really good things about what it caught in terms of these longunning tasks is in inbox triage, it ignored a prompt injection,” Vo noted. The model processed the email, extracted the necessary action item, and bypassed the injected command completely.
Beyond security filters, the model showed pattern awareness across messy raw datasets. “It also can kind of like reason with the context of a lot of research. And so, it found that in some research that I had to do, 41 out of the 44 Confluence tickets came from one customer,” Vo said. It also flagged customer requests that violated internal company policies, catching edge cases that human reviewers frequently miss when skimming large backlogs.
The Dual-Model Development Stack
Instead of standardizing on a single AI provider, Vo restructured her engineering process around a split model setup. She runs OpenAI Codex for rapid execution while deploying Opus 5.5 as a validator.
Splitting workloads across two independent model families prevents blind spots. If Codex writes a database migration, Opus 5.5 reviews the logic, queries, and security surface before anything hits production.
What to Do With This
Pick an existing agent prompt this week and test it against an adversarial data sample. Insert a hidden override instruction into a raw customer email or JSON payload to verify whether your current model follows your system instructions or succumbs to the injection. If your pipeline fails, route high-risk multi-step parsing through Opus 5.5 and use parallel models to cross-verify code output.