Key Takeaways
- Working with AI research teams requires a different toolkit than standard engineering, focusing on user failure modes instead of sprint tickets.
- OpenAI product leader Tara Seshan emphasizes that sample readings of real user session logs reveal exact friction points where models fail.
- Writing rigorous evaluations (evals) is the most valuable technical skill an AI product manager can develop.
- Prototyping successful prompts and tool skills gives post-training teams a reproducible baseline to train directly into base models.
- Product managers can drive base model improvements through Seshan's 4-Step AI PM to Research Feedback Loop.
The Seshan's 4-Step AI PM to Research Feedback Loop
Tara Seshan discovered early at OpenAI that managing AI products does not match standard software delivery. “My take coming in working with research was, oh, working with research is different than working with engineering,” Seshan explained. “Your value to research is knowing specifically what the users want and what their use cases are.”
To bridge the gap between user intent and research priorities, Seshan developed a four-stage process:
- Step 1: Conduct Sample Readings and Isolate Failures: Come with very specific use cases, clear user goals, and analyze exact session transcripts to understand why the model failed to achieve the desired outcome.
- Step 2: Write Rigorous Model Evaluations (Evals): Author evals to objectively measure and benchmark model performance against the identified failure modes.
- Step 3: Prototype via Prompts and Skills: Demonstrate proof-of-concept by proving that specific prompt designs and tool skills successfully accomplish the task with current model capabilities.
- Step 4: Transfer Prototypes to Post-Training: Take the successful prompts, skills, and eval datasets directly to the post-training team so the capability can be fine-tuned and trained into the base model from the start.
When This Works (and When It Doesn't)
This loop works when product teams collaborate directly with foundation model researchers to translate observed user pain points into native base model improvements. If your team trains or fine-tunes custom models, handing researchers concrete eval sets and working prompt scaffolds gives them clear optimization targets. As Seshan notes: “Learning how to write good evals and writing evals as much as possible is the most important thing.”
The framework fails when applied to off-the-shelf API wrappers without access to model weights or post-training pipelines. If your startup relies entirely on third-party commercial APIs, research cannot retrain base model behavior for your specific vertical. In that environment, your prompt-and-skill prototype remains your permanent production architecture rather than an intermediate handoff to post-training.
What to Do With This
Assume your customer support agent fails when users switch topics mid-conversation. Do not file a vague bug ticket stating that the model gets confused.
First, pull twenty failed conversation transcripts and document the exact user prompt where context broke. Second, turn those twenty failures into a benchmark eval with binary pass-fail criteria for context retention. Third, build a system prompt paired with a memory retrieval skill that achieves a ninety percent pass rate on that benchmark. Finally, take that dataset and prompt prototype to your ML team to bake that multi-turn tracking directly into your fine-tuned model checkpoint.