Key Takeaways

  • ElevenLabs scaled past $450 million in ARR and an $11 billion valuation by committing to cascaded voice pipelines instead of unified speech-to-speech models.
  • Cascaded voice architecture chains three distinct modules: speech-to-text (STT) transcription, reasoning via a large language model (LLM), and text-to-speech (TTS) audio generation.
  • Direct speech-to-speech models cut latency by removing intermediate text conversion, but they rely on smaller parameter architectures that struggle with complex reasoning and exhibit higher hallucination rates.
  • Enterprise voice agents require deterministic tool execution, identity verification, and database lookups, capabilities that black-box speech-to-speech models cannot inspect or verify mid-flight.

The Disagreement

AI research labs argue that direct speech-to-speech models represent the clear future of conversational interfaces. By training a model to ingest audio tokens and emit audio tokens directly, you eliminate the latency overhead of speech-to-text and text-to-speech conversions. You also preserve vocal inflections, laughter, accents, and emotional tone that get flattened when voice is converted into plain text.

ElevenLabs co-founder Mati Staniszewski takes the opposite side. Staniszewski argues that end-to-end models sacrifice the single thing business customers actually care about: deterministic control.

“As you said, our approach, as you think about voice agent, conversational agent, is effectively a cascaded approach,” Staniszewski explained to Stripe co-founder John Collison. “You use transcription or speech-to-text, LLM, text-to-speech, and orchestrates all of that together.”

When Collison pressed on why labs chase direct models, Staniszewski pointed out the hidden cost of collapsing the pipeline into one model. “It's quicker, but on the flip side, you lose reliability. You lose like all visibility into the parts of the pipeline.”

Collison summarized the practical outcome bluntly: “Do you observe that speech-to-speech models think differently than cascaded models? It sounds like they're dumber.”

Who's Right (and When They're Wrong)

Staniszewski is right for any system that executes real-world actions, but end-to-end models will still capture pure entertainment use cases.

If you are building an open-ended conversational companion, language tutor, or voice game, speech-to-speech models shine. In those contexts, a five-hundred-millisecond drop in latency and natural emotional pacing matter far more than factual accuracy. A minor hallucination in a friendly chat ruins nothing.

For enterprise workflows, speech-to-speech models fall apart. A voice agent answering customer support calls for a bank or an airline cannot simply guess what to say next. It must transcribe the customer request, authenticate the caller against a database, verify permissions, call an API to fetch booking details, and inspect the intermediate text before speaking the resolution aloud.

“As we work with a lot of the businesses and enterprises, they will need that visibility into what happens,” Staniszewski noted. “They will want to execute certain tasks on top of that.”

In a cascaded stack, engineers can inspect the prompt between STT and LLM, enforce strict guardrails on the output text before synthesis, and swap in faster or larger models for specific tasks without retraining the entire voice engine.

What to Do With This

If you are building an AI voice product this week, map every user interaction to a strict requirement test. If your agent executes database writes, processes financial transactions, or triggers API calls, build on a modular cascaded pipeline with explicit text inspection layers. Reserve end-to-end speech models exclusively for low-stakes conversational flows where sub-second latency matters more than execution accuracy.