Key Takeaways
- Mati Staniszewski and Piotr Dąbkowski started ElevenLabs in 2021 to fix foreign movie dubbing, motivated by Poland's single-voice narration style.
- Dubbing requires three distinct steps: transcription, translation, and speech regeneration. In 2021, research for each piece produced robotic, unusable results.
- Instead of building a flawed end-to-end translation pipeline, ElevenLabs focused entirely on text-to-speech research to make synthesized voices sound human.
- Early feedback from creators revealed immediate demand for narrating scripts and editing audio, proving that fixing text-to-speech solved a live commercial problem before multi-language dubbing was ready.
- Building in 2022 allowed ElevenLabs to focus on AI voice research while tech attention chased crypto and metaverse trends.
The Monotone Voice in Polish Television
If you grew up watching foreign movies on Polish television, you never heard distinct actors speak in translated releases. You heard a single reader, known as a lektor, speaking flatly over the original audio track.
Staniszewski experienced this frustration firsthand. “The actual trigger point comes from where you're from, from Poland. Very peculiar thing. If you watch a movie in Polish, all the voices, whether it's a male voice or whether it's a female voice, get narrated with one single character,” Staniszewski said. “So you have one voice narrating the whole movie. All the emotional intonation disappears.”
In 2021, Staniszewski and his co-founder Piotr Dąbkowski set out to replace that outdated workflow with automated dubbing software. They wanted characters in translated movies to keep their original emotional weight, cadence, and distinct identity across every language.
Deconstructing the Dubbing Chain
When Staniszewski and Dąbkowski mapped the requirements for full automated dubbing, they found a chain of three separate technical steps: transcribing source audio into text, translating the text into the target language, and regenerating the translated text into realistic speech.
At the time, attempting to build all three steps at once compounded errors. “Initially, it was dubbing, and then, as we started diving into what we need to do to solve dubbing, realized there are three steps in dubbing process,” Staniszewski explained. “There is transcription step, then it's translation step to another language, and then you need to regenerate that in another language. But the research that existed at the time for each of those steps wasn't very good.”
If the final speech generation step sounded robotic, the entire dubbing pipeline collapsed. A perfect translation still felt unwatchable if the generated voice lacked natural cadence and human emotion.
“And that for us, it was like, okay, before we can solve dubbing, let's solve the research component to generate speech and make it sound great,” Staniszewski said. The founders paused their broad dubbing vision to focus on pure speech synthesis.
Perfecting Single-Language Narration First
To build a working product, ElevenLabs cut scope. Instead of handling multiple languages and synchronization, they focused on letting creators generate expressive single-language speech from written text.
This decision gave them an immediate product that content creators and authors could use to narrate books, produce video voiceovers, and edit audio scripts. It also gave them focus while the rest of the market looked elsewhere. “It was still a year when the topics of the day were crypto and metaverse. So 2022 was still a year where everybody was obsessed about those two. So it was a perfect time because we could actually focus and build a lot on the AI side.”
By the time multi-language dubbing became technically viable, ElevenLabs had already established the standard for synthetic voice quality.
What to Do With This
Audit your product roadmap and identify the weakest link in your core user workflow. If your end-to-end product relies on three technical steps and one step delivers poor quality, stop building the full chain. Strip away the surrounding steps, isolate the broken technical component, and turn that single solved capability into your initial standalone product.