Key Takeaways

  • Early ElevenLabs conversational agents flopped because their speech was unnaturally articulate; adding human disfluencies like pauses, "ums," and "uhs" caused user performance and acceptance to skyrocket.
  • The company has restored voices for more than 10,000 patients suffering from ALS or throat cancer, enabling people like former NFL player Tim Green to continue hosting podcasts.
  • Voice carries emotional identity through intonation and timing that text prompts alone cannot recreate.
  • Staniszewski keeps ElevenLabs strictly focused on audio research, applying small, flat engineering teams inspired by his time at Palantir.

The Flaw of Artificial Perfection

Engineers naturally optimize for precision. When ElevenLabs built its earliest voice agents, the team tried to eliminate every stutter, pause, and verbal glitch. They engineered synthetic speakers that pronounced every syllable with pristine clarity.

Users hated it. The voices felt sterile, synthetic, and alienating. People did not trust what they heard because real people do not speak in edited paragraphs.

Staniszewski observed a clear pattern in how humans process audio. As he put it: “Voice carries so many other dimensions than text. Like text, of course, you imagine, you interpret, but it doesn't have the emotion, it doesn't have the intonation, it doesn't have the imperfections, it doesn't have the pauses.”

To fix the problem, the team reversed course. They deliberately injected the flaws of real speech back into the models.

“Initially, we are trying to create a perfect voice agent that doesn't do any imperfections. It didn't sound human,” Staniszewski explained. “And then of course the obvious thing, and that was the clearest, it's like the 'Ums,' the 'Uhs,' the pauses and suddenly, the performance of working with that voice agent skyrocket.”

Realism in conversational interfaces does not come from mathematical optimization. It comes from mirroring human friction. When an agent pauses to gather a thought or drops a conversational filler, the listener relaxes. The brain stops analyzing the machine and starts listening to the message.

Restoring Identity Beyond Text

Staniszewski grew up in Poland frustrated by flat, single-voice movie dubbing where one monotonous narrator read dialogue over American actors. That irritation drove him to explore frontier voice synthesis, but the technology quickly outgrew media production.

Voice is tied directly to biological identity. As Staniszewski noted, “The moment you hear someone's voice, you recognize it if you know someone.” Losing vocal ability strips away a person's immediate social presence.

ElevenLabs put this into production by working directly with individuals facing degenerative diseases. The company has restored synthetic voices for over 10,000 people who lost their ability to speak due to ALS or throat cancer. These recovered voices have allowed users to give wedding vows, speak before Congress, and continue professional careers, such as former NFL player Tim Green hosting his podcast.

For Staniszewski, this application proves the real test for artificial intelligence: “AI really needs to work for the people. It needs to amplify human potential rather than replace it.”

When software anchors directly to human identity instead of abstract automation, users form an immediate, visceral connection to the product.

What to Do With This

Audit your user-facing AI prompts, agent scripts, or copy this week. Strip out the sanitized corporate language, unbroken cadences, and overly polished summaries. Introduce natural pacing, shorter phrasing, and conversational beats that match how your users actually talk to their peers.