Key Takeaways

  • Affirm rolled out Arc, a transformer-based attention underwriting model that doubled the performance gains of their previous gradient-boosted tree models.
  • Levchin dismissed the threat of standalone agentic payment startups, noting that established fintechs with capital-market access can adopt agentic workflows faster than newcomers can build balance sheets.
  • In January, Affirm carved out an isolated 12-person engineering team to benchmark developer tools and remove ad-hoc experimentation from the rest of the company.
  • The prescriptive approach drove a 10x jump in human-machine hybrid code production while reducing fully loaded cost per pull request by 30% across engineer salaries and AWS compute.
  • Levchin formalizes this strategy as Affirm's Centralized Developer Productivity Model.

The Centralized Developer Productivity Model

Step 1: Dedicate an Isolated Tooling Team

Carve out a dedicated unit of senior engineers (e.g., 12 people) whose sole mandate is benchmarking coding harnesses, models, and workflows rather than product development.

Step 2: Eliminate the Tyranny of Choice

Stop broad ad-hoc developer experimentation across the entire organization. Provide a curated, prescriptive menu of vetted developer environments and AI coding tools (e.g., Cursor, Claude).

Step 3: Enable Frictionless Model/Harness Swapping

Abstract underlying tools so that when state-of-the-art changes, the core engineering organization can migrate to the superior model or harness seamlessly without disruptive tooling overhauls.

Step 4: Track Fully Loaded Unit Cost per Output

Measure engineering productivity using standard business metrics, specifically tracking fully loaded cost per pull request (including engineer compensation and cloud/compute costs).

When This Works (and When It Doesn't)

This setup works when an engineering team passes 50 to 100 people and developers start spending hours testing every new coding assistant on Twitter. When individual engineers run their own experiments, company codebases fragment across conflicting plugins, security policies break, and compute bills spike without measurable output gains. A centralized team absorbs the evaluation overhead and delivers a clean setup to everyone else.

It fails at early-stage startups with fewer than 20 engineers. If you pull two senior developers off product delivery to run harness benchmarks when you do not yet have product-market fit, you burn runway on premature optimization. Small teams should pick one standard stack, enforce it lightly, and get back to shipping customer features.

What to Do With This

If you manage an engineering organization spending real money on AI tooling, run an audit this Friday:

1. Calculate your true baseline unit cost. Take your total monthly engineering spend (base salaries, contractor fees, cloud compute, and AI seat licenses) and divide it by the number of merged pull requests over the last 30 days.

2. Stop allowing every developer to expense a different AI extension. Appoint one senior engineer for two weeks to test the top two setups against your actual repository.

3. Standardize on the winning tool, set up a shared abstraction layer for API keys and models, and mandate that every engineer use the approved workflow.

4. Measure the fully loaded cost per pull request 60 days later to confirm whether your output per dollar actually improved.