Key Takeaways
- Every CEO Dan Shipper separates exploratory AI lab experiments from core roadmap execution to prevent experimental chaos from breaking stable releases.
- Shipper pairs rapid prototype builders ("pirates") with systems engineers ("architects") on small two-person teams to turn rough scripts into durable software.
- Prototype retention is tested internally first: if team members do not repeatedly use a prototype on their own without prompting, the prototype is killed.
- New model capabilities must pass a 30-day durability test to verify they remain 10x better than existing baselines after the initial novelty fades.
- Shipper tracks ideas through The 4-Stage AI Research Pipeline and Promotion Criteria during weekly all-hands meetings in Notion.
The 4-Stage AI Research Pipeline and Promotion Criteria
Stage 1: Lab-Only Experiments
Run rapid, disposable parallel experiments on new model capabilities without roadmaps or polish, expecting to discard roughly 90% of output.
Stage 2: Internal Dogfooding & Team Adoption
Test surviving prototypes on real internal workflows. Measure whether colleagues adopt the tool repeatedly on their own without prompting.
Stage 3: Early Customer Validation & Architect Refactor
Pair a pirate prototype with an architect engineer to build metrics dashboards, evaluate unit economics, and test with VIP early adopters.
Stage 4: Core Product Merge & Scale
Hand off proven features to the main product roadmap using three criteria: (1) internal retention/repeat use, (2) 10x improvement over existing baselines 30 days later, and (3) scalable affordability.
When This Works (and When It Doesn't)
Shipper built this pipeline to solve a specific headache: how to continually feed new AI breakthroughs into core products without disrupting roadmap stability. As Shipper explains, “The thing you need to do is make a research pipeline where ideas start on the left and they start out as lab only and there are lots of little experiments and then you move them progressively from the left to the right and at some point you have a handoff with your product team where they start to get incorporated into the product.”
This setup works when underlying AI models change weekly and your engineering team cannot afford to refactor production infrastructure for every new release. It gives rapid builders the freedom to hack without breaking uptime SLAs. When your testing pipeline functions well, unexpected model releases stop causing internal panic: “The way that you know this is working is you will welcome moments when new models drop and be excited about them instead of dreading them.”
Where this system fails is in enterprise domains with strict compliance, SOC2 audits, or rigid security controls. In those environments, ungoverned lab scripts touching real customer data will trigger legal blocks. It also breaks if you lack architects who can translate messy experimental code into maintainable systems.
What to Do With This
If you run a product team, set up a simple board in Notion this week with four columns representing each pipeline stage. Assign one engineer to act as your pirate, giving them 48 hours to build throwaway prototypes using a new model release.
Next, run a strict filter before any prototype touches production code. Shipper stresses that promotions must be deliberate: “Another really important thing is defining clear decision criteria for moving an idea through the pipeline and really making it an event when you do it.” Review the pipeline during your weekly all-hands. Ask Shipper's core question: “Is it 10x better than what currently exists? And this is actually a really interesting filter because what is new feels very exciting, but the question is when it is a month later, is it actually any better?” If your own team stopped using the prototype after week one, kill it immediately and clear the board for the next experiment.