Key Takeaways

  • Mike Krieger built Anthropic's first computer-use prototype in 2024, but the models were too weak to support reliable automation or user education.
  • Instead of abandoning failed experiments, Anthropic converted user tasks into an automated eval suite and ran it against every new model checkpoint.
  • Claude 3.7 surfaced as a breakthrough when it suddenly completed parked tasks, turning a dormant codebase into an active product.
  • Dan Shipper calls the risk of discarding failed experiments capability blindness: assuming a task is impossible when a model release three months later solves it easily.
  • Teams building ahead of current foundation models can systematically bridge research and product roadmaps using the Anthropic Labs Frontier Eval Parking Method.

The Anthropic Labs Frontier Eval Parking Method

  • Step 1: Build Ahead of Current Capabilities: Prototype ambitious agent workflows and product concepts early to establish the target interaction model, even if current models fail execution.
  • Step 2: Externalize Failures as Evals: When a prototype fails because models are not ready, convert the failure modes and user tasks into an automated evaluation harness and connect findings with research teams.
  • Step 3: Park and Auto-Benchmark New Model Releases: Park the product build and continuously run the benchmark harness against every subsequent model checkpoint without active manual rebuilds.
  • Step 4: Detect Capability Leaps and Unpark: Inspect eval transcripts when automated pass rates suddenly jump, validating that the new model has crossed the threshold to unpark and productize the feature.

When This Works (and When It Doesn't)

This system applies when building software that fails due to model limitations rather than flawed product thesis, weak demand, or broken interface design. It stops teams from abandoning viable features right before the underlying reasoning capabilities arrive. As Krieger explained: “it was so bad because the models just weren't there yet... we parked it and then we would basically just try it with every new model. And we actually, the way we would try it is we actually just had it running in an eval harness.”

It breaks down when the product hypothesis itself is flawed. If users do not want an autonomous agent touching their spreadsheets, a smarter foundation model will not fix the adoption problem. Parking bad concepts in test suites only generates automated noise. Reserve this practice for workflows where the interaction model is clear and execution reliability is the sole blocker.

What to Do With This

Take an agent workflow your team shelved over the past six months because output quality was unreliable. Revive it as an automated benchmark this week:

  • Step 1: Pull your team's prototype for autonomous invoice reconciliation. You already have the target interaction flow and expected outputs defined.
  • Step 2: Convert 20 actual failure cases from your test logs into a deterministic test script, capturing the prompt, input files, and expected state changes. Share these failure transcripts with your research contacts.
  • Step 3: Add that test script to your CI pipeline to run automatically against every new model release or checkpoint update without manual engineer rebuilds.
  • Step 4: When a new reasoning model crosses an 85% success rate on the test suite, inspect the execution transcript, unpark the repository, and ship the beta to users.