Key Takeaways
- Claire Vo processed 4,500 YouTube comments with TypeSafe AI's decision model Jev, replacing expensive frontier LLM calls with fast, typed classification primitives.
- Vo categorized sentiment into positive, negative, or neutral using choice primitives, and extracted topic recommendations using boolean flags.
- Live search across thousands of text records required architectural batching: running fast score models across comment batches and filtering for top scores to achieve sub-second speeds.
- Small decision models solve classification and search pipelines at a fraction of the latency and cost of generative models.
The Method
Founders often default to giant frontier models when they want to analyze customer text. If you want to sort 4,500 records, sending full prompts to a general-purpose model is slow, expensive, and fragile. Vo tackled this by treating classification as a set of small, typed operations.
First, Vo pulled raw audience feedback using the YouTube Data API enabled in Google Cloud Console. “And there's about 4,500 comments and I just wanted to know like are they good, are they bad, are they happy, are they sad? And I also wanted to know if you all had any episode ideas for me,” Vo explained.
Second, she structured the analysis into discrete primitives. Instead of asking a model to write a freeform summary, she assigned typed tasks. One task was sentiment categorization: “You just have to enable it in Google console. And then I said pull all the comments and categorize them into positive, negative or neutral. So that would be a choice, right? Or a score probably.” A second pass applied a boolean test to identify future episode topics. Vo then aggregated those structured outputs into an interactive dashboard.
Third, she engineered sub-second live search over the comments. If a user searched for a term like "screen share," the system needed to scan the dataset and return matches instantly. Vo avoided slow sequential passes by batching the data. As Vo noted, "just for like behind the scenes, you have to like batch the results and then score them and then push the high scores up to get that kind of performance."
Where This Breaks Down
Decision models run fast because they output narrow types like scores, choices, and booleans. That architecture works when the question is clear: is this comment positive, and does it suggest a topic?
It breaks down when you need complex synthesis or cross-document reasoning. If an audience member writes a multi-paragraph technical critique that mixes valid product feedback with edge-case feature requests, a simple boolean flag loses the context. You still need frontier models to synthesize themes and draft responses. Using a decision model for semantic search also requires careful batch sizing; make batches too large and latency spikes, make them too small and network overhead kills your live search.
What to Do With This
Export your last 1,000 customer support tickets or app reviews into a CSV file tomorrow morning. Set up two typed classification checks: a sentiment score and a boolean flag for feature requests. Run them through a dedicated decision model pipeline in batches of 50 to generate an automated product signal board without paying frontier LLM token rates.