Key Takeaways

  • Anthropic's Thariq Shihipar predicts context files like CLAUDE.md and AGENTS.md will disappear as frontier models improve.
  • Storing historical bug patches in markdown files overconstrains newer models that no longer share those failure modes.
  • Frontier reasoning models become Pareto dominant over smaller models by burning fewer total tokens on verification.
  • Effort allocation should match domain stakes: set reasoning effort to high or max for security audits and code reviews, but lower for general software engineering.

The Problem with Context Files

Engineers love writing context files. You build an agent harness, watch Claude trip over a specific library bug or hallucinate a syntax rule, and log a fix inside CLAUDE.md or AGENTS.md. Over six months, that file turns into a sprawling graveyard of edge cases.

Shihipar says this practice hurts newer models instead of helping them. “As the models get better and better, the floor of how they accomplish the simpler task is better,” Shihipar explained. “And so I do think in the limit CLAUDE.md goes away.”

When you upgrade your model version, your historical log turns into noise. If an earlier model version struggled with a specific framework quirks, writing negative constraints into context tells the new model to avoid paths it could otherwise handle cleanly. Shihipar pointed out the trap: “If you've added a bunch of failure modes... maybe Fable 5 had this failure mode that Fable 5.1 doesn't. And if you keep this context running log of a bunch of different failure modes, they will probably overconstrain Claude.”

Frontier Models and Pareto Dominance

Common intuition says small models are cheaper for simple coding tasks while giant reasoning models belong on hard architectural problems. Shihipar argues the economics are flipping.

“The frontier models will be Pareto dominant over almost everything,” Shihipar stated. “Increasingly it's just going to be that the smart model is going to be able to do the simple task for less tokens than the other models, basically because of verification.”

A smaller or weaker model often generates flawed code, catches errors late, runs redundant test loops, and generates hundreds of extra tokens trying to fix its own mistakes. A smarter frontier model gets the code right on the first attempt. Because it skips repeated verification cycles, the smarter model uses fewer total tokens, wiping out the per-token price discount of smaller architectures.

How to Allocate Model Effort

Reasoning models allow developers to adjust thinking effort. Shihipar warns against applying uniform effort settings across your entire engineering workflow. Effort must track task complexity and error cost.

“Effort scales with basically the complexity of the task,” Shihipar said. “So for security, effort gets way more results. High effort versus low effort changes the eval. But for software engineering it doesn't change it a huge amount because effort is mostly spent on the verification and the edge case testing.”

His rule of thumb for engineering teams is simple: “My rough distribution is code review and security should be high or max basically, and software engineering settings per domain.”

In standard feature development, high reasoning effort yields diminishing returns because fast test runners handle validation. In code review and security analysis, missing a single vulnerability costs thousands of dollars, making maximum reasoning effort the only logical choice.

What to Do With This

Open your repository's CLAUDE.md or agent system prompts today and delete every negative instruction added more than two model versions ago. Split your agent configs into two profiles: lock routine code generation to medium reasoning effort, and set your pull request review and security scanners to maximum effort.