Key Takeaways

  • Claude Code and Codex store complete user turns and session history directly on your local hard drive.
  • Claire Vo parsed her local session files and discovered pure engineering dropped from nearly 100% of her activity in January to under 40% in September.
  • High-speed decision models like Jev classify thousands of local user prompts into categories like coding, agent management, and media production for pennies.
  • Auditing your session history reveals hidden shifts in your daily workflow before you notice them manually.

The Method

Most software engineers guess how they spend their time. They assume they write code for eight hours. In reality, their actual work shifts rapidly as new AI tools enter the stack.

Claire Vo, founder of TypeSafe AI, tracked this shift directly from her desktop. Tools like Claude Code and Codex store their logs locally. Every prompt, code change, and user turn sits in local session files. Vo ran TypeSafe AI's fast decision model, Jev, directly over these raw session logs to categorize each turn by task type.

“All of your Claude Code and Codex sessions are stored locally,” Vo said. “You can actually run this analysis on everything stored on your local machine.”

Because small decision models classify structured text rapidly at near-zero cost, you do not need expensive frontier model calls to audit months of logs. Vo grouped her sessions by day and user turn. The results surprised her.

“For grouping by user turns in January I was almost exclusively doing engineering tasks in Claude Code and Codex,” Vo noted. “And then as you kind of come into this new world now in September less than 40% it looks like of my tasks are actually engineering tasks. I'm doing a lot more work with agents which makes a lot of sense. And then I'm doing a lot more like publishing media.”

The remaining time spread across agent orchestration, client delivery, and media generation. The audit showed her exactly where manual coding stopped and workflow management took over.

Where This Breaks Down

This audit only tracks what happens inside local terminal and editor sessions. It misses whiteboarding, team syncs, PR reviews in browser interfaces, and offline architecture planning. If you spend three hours debugging a design flaw on paper before writing one prompt, the log only records the single prompt turn.

Classification accuracy also depends on your taxonomy. If your prompt classifier uses vague categories, agent workflow tasks get lumped together with standard code generation. You need explicit labels: pure code authoring, test execution, agent orchestration, and documentation.

What to Do With This

Locate your local session directory for Claude Code or Codex this week. Write a script to extract every user turn into a JSON list. Run a lightweight classification model across the extracted turns with four distinct categories: code creation, agent debugging, documentation, and operational delivery. Chart the distribution across the last six months to see where your personal bottlenecks are moving.