How Descript engineers multilingual video dubbing at scale

Descript redesigned its translation pipeline using OpenAI reasoning models to optimize for semantic fidelity and duration adherence, resulting in a 15% increase in dubbed exports.
Results
43
Percentage point improvement in duration adherence with OpenAI
Results
15%
Increase in dubbed exports post-rollout
Descript(opens in a new window) is an AI-native video editor built around a simple idea: if you can edit text, you should be able to edit video. Since Descript’s early days, AI has powered every aspect of the product: transcription, editing, audio cleanup, and increasingly complex creative workflows. They’ve built on OpenAI for years, using Whisper for transcription and GPT series models inside their co-editor Underlord.
Translation quickly emerged as a high-impact use case. Traditionally, translating video has been slow and expensive, requiring language experts to manage projects, produce rote translations, handle quality control, and generate corresponding audio. LLMs dramatically compress that workflow, making high-quality translation at scale possible.
Captions and dubbing both require semantic fidelity: the translation must preserve the original meaning. But duration adherence plays a different role in each. For captions, it's a nice-to-have. For dubbing, it's critical, because if translated speech runs too long or too short, it will sound unnatural even if the meaning is correct.
To address this, Descript redesigned its translation pipeline using OpenAI reasoning models to optimize for semantic fidelity and duration adherence during generation, not after. In the first 30 days after rollout, exports of translated videos with dubbing increased 15%, and duration adherence improved by 13 to 43 percentage points, depending on the language.
“Dubbing is an increasingly popular use case for Descript, so we’re building ways to do it in batch for companies that want to translate and lip-sync entire libraries,” said Laura Burkhauser, CEO.
Translation was one of Descript’s earliest and most requested features. They started with captions-only translation, which worked well—but many users wanted to go further and have spoken audio (dubbing) in the target language.
However, one issue kept surfacing: dubbed audio didn’t always sound right. “Probably the number one complaint we heard was that the pace of the speech was unnatural in the translated language,” said Aleks Mistratov, Head of AI Product at Descript.
The problem came down to the fact that different languages take different amounts of time to express the same idea. Descript observed, for instance, that on average German is a “longer” language than English. To fit into fixed video segments, translated speech often had to be artificially sped up or slowed down. “You’d end up with something that sounded like chipmunks, or a sleepy giant,” Mistratov explained.
Source: OpenAI News















