Key Takeaways

  • Text models passed the conversational Turing test years ago, but voice models struggle with the physics of real-time speech orchestration.
  • ElevenLabs scaled past $450 million in annual recurring revenue and an $11 billion valuation by solving voice synthesis, yet full conversational parity remains open.
  • Narrow enterprise workflows like customer service authenticate users and pass conversational tests today, while open domains like gaming remain years away.
  • Database lookups and authentication create latency spikes that break conversational illusions faster than model hallucinations.

The Disagreement

John Collison believes consumer voice tech remains fundamentally broken. “There was no way I could get my phone to read me something, which seemed like a fairly basic feature,” Collison noted. “All cars advertise voice control, and yet, it sucks.” His core argument is simple: text models reached human-level output, but voice agents still trip over basic speech mechanics.

“That's the simpler way of saying what I'm saying is that we have passed the Turing test with text LLMs a long time ago. We're actually nowhere near that on voice LLMs,” Collison argued.

Mati Staniszewski sees a clear split between narrow workflows and open-ended conversation. He agrees with the broader diagnosis on general voice agents, but points out that vertical tasks are already crossing the human threshold.

“I agree with the claim that this orchestration side has not passed a true conversational agent Turing test, where it behaves as you would expect from another person,” Staniszewski said. Yet in focused domains like standard customer support calls, structured intent allows systems to route queries and speak naturally right now.

Who Is Right (and When They Are Wrong)

Both founders are right, but they are looking at two different engineering problems.

Collison is evaluating open-ended voice agents. In an unconstrained setting, humans manage constant micro-cues: mid-sentence pauses, interjections, tone shifts, and instant interruption recovery. Traditional software cascades stack speech-to-text, an LLM call, and text-to-speech sequentially. Each step adds 200 to 500 milliseconds of latency. By the time the system decides whether you finished your thought or merely took a breath, the conversational flow is dead.

Staniszewski is evaluating constrained business workflows. A customer calling a bank has a bounded set of intents. The system does not need to simulate human wit; it needs accurate retrieval and low latency.

The breaking point happens the moment the agent needs external context. As Staniszewski pointed out: “If it's a conversational use case, pretty simple, you can root the agent to speak with it. But if you need to authenticate, if you need to pull additional information from the database, what do you do? How do you handle that graciously? That's where it gets tricky.”

If your agent takes 1.5 seconds to query a Postgres database while a customer waits, silence feels like a dropped call. Human agents bridge that gap with filler speech ("Let me pull up your account right now"). Voice models that do not orchestrate filler phrases during async operations will instantly fail the voice Turing test.

What to Do With This

If you are building voice features this week, stop trying to build a general companion. Restrict the agent's operating scope to tasks with three or fewer user intents. Insert hardcoded filler phrases or auditory feedback during database fetches so your latency floor never creates dead air on the line.