Most teams discover the limits of dubbing the same way: a winning ad in one market gets translated into six others, the spend follows, and the results come back flat everywhere except the original language. The conclusion drawn is usually that the other markets are worse, or that the offer does not travel. Occasionally that is true. Far more often the creative was never localized at all - it was translated, which is a much smaller thing, and the audience noticed before they could explain why.
It is worth being precise about what "localized" means for a video ad, because there are three genuinely different products hiding under the word, and they are not three price points on the same ladder.
One brief, every market
siriusly.ai generates a separate reel per language - native speaker, native cadence, market-appropriate creator - instead of dubbing one take into many.
Try it freeThree Ways to Ship an Ad in Another Language
| Subtitle | Dub / lip-sync | Native generation | |
|---|---|---|---|
| Audio | Original language | Translated, timed to original | Written and spoken natively |
| Performance | Original | Original body, new mouth | Generated for this script |
| Casting | Fixed | Fixed | Per market |
| Length | Fixed | Forced to match original | Free |
| Cost per market | Near zero | Low | Cost of one generation |
Subtitling is the honest option and it is fine for informational content. For a UGC-style ad it costs you the sound-on audience and, more importantly, it makes the ad visibly foreign - which is the exact opposite of the format's premise that this is a person like you, in a room like yours.
Dubbing is where most budget goes, and where most of the loss happens. The rest of this post is about why.
What Actually Breaks in a Dub
The performance belongs to the original language
Speech rhythm is not decoration on top of meaning; it is part of how a sentence is understood. English delivers information in a fairly flat stress-timed pattern with emphasis carried by pitch. Spanish and Italian are syllable-timed and faster. Japanese carries emphasis through particles and sentence-final structure that lands in a completely different place in the sentence. German routinely holds the verb until the end, which means the moment of comprehension - the beat a good creator naturally reacts on - is at a different point in the clause than the English original.
When you dub, the body language, the eyebrow raises, the head nods and the hand gestures all stay pinned to the English comprehension beats. The new audio lands somewhere else. Nothing is visibly broken frame by frame, but the performance is emphasizing the wrong words, and the viewer reads it as a person who is not quite present in what they are saying.
Length is forced
Translated text is rarely the same length as its source. German and Finnish typically run 20 to 35 percent longer than English; Spanish and French land 15 to 25 percent longer; Japanese and Chinese are often shorter in characters but not always in spoken duration. A dub has to fit the original timing regardless, which leaves two options and both are bad: speak faster than anyone naturally would, or cut content out of the script.
Dubbed: the German script has to say the same thing in the same 28 seconds, so it is either rushed to the point of sounding like a disclaimer or trimmed until the demonstration beat is gone.
Native: the German script is written to be spoken in German, at a German pace, and the reel is simply 33 seconds instead of 28. Nothing is cut and nothing is rushed.
Casting is frozen at the original market
This is the largest and least discussed cost. A UGC ad works because the viewer recognizes the speaker as someone plausibly like them. Dubbing keeps one face across every market, which means at most one market gets a creator who looks and sounds local, and everyone else gets a foreigner speaking their language unusually well. In markets where the product's credibility depends on the speaker being a peer, that is a real conversion cost, not an aesthetic quibble.
The mouth region gives it away
Lip-sync models have improved enormously and the best of them survive a casual watch. What they still struggle with is the interaction between a re-synthesized mouth and everything around it: facial hair, teeth at wide vowels, head rotation past a certain angle, and any frame where a hand passes near the face. On a phone screen at arm's length, most viewers cannot name the artifact. They can still tell that something is synthetic, and in an ad format built entirely on looking un-produced, "something is synthetic" is the one impression you cannot afford.
Where Native Generation Actually Wins
Generating each market separately is not just a quality upgrade to dubbing - it removes the constraint that caused the problems. Because nothing is pinned to an original take:
- The script is written, not translated. Idioms become local idioms. Price framing uses local currency and local price psychology. A claim that is legally routine in one market can be softened where it is regulated in another.
- The creator is cast per market. Ethnicity, age range, wardrobe and room can all match the audience rather than the source video.
- Duration is free. Each language runs as long as it naturally needs.
- The voice is native. A speaker with a plausible regional accent, not an accent transferred from the source language.
- Nothing degrades. Every market is a first-generation render rather than a modification of one.
The Detail Everyone Underestimates: Numbers
Numbers, units, prices and percentages are the most reliable source of embarrassing localization failures, because a script is written to be read and a voice model reads it literally. The failures are specific and they repeat:
| In the script | What a naive read produces | What a native speaker says |
|---|---|---|
| 24g | "twenty-four gee" | "twenty-four grams" |
| 0,0% (German) | digits left unvoiced or read as English | "null komma null Prozent" |
| CIA | attempted as a word | "C-I-A", spelled out |
| NASA | spelled out letter by letter | "NASA", as a word |
| CapCut | one unpronounceable blob | "Cap Cut" |
| $19.99 | currency symbol dropped | local currency, local decimal convention |
Every one of those is a small thing that instantly marks an ad as machine-made. The fix is not clever prompting after the fact; it is normalizing the script into spoken form before it reaches the voice or video model, per language, including which acronyms are spelled and which are pronounced.
Captions Have to Be Localized Too
Burned-in captions are part of the creative, not a subtitle track bolted on, and they localize badly if treated as an afterthought:
- Line length. A caption style tuned for English words breaks awkwardly on German compounds. Either allow a smaller font per language or set the wrap width per language.
- Script coverage. Arabic, Hindi, Japanese, Korean, Chinese and Russian need a font that actually has those glyphs. A missing glyph renders as a box, and in a burned-in caption that box is permanent.
- Direction. Arabic runs right to left, which affects both alignment and where a word-by-word reveal should start.
- Timing. Word timings must come from the localized audio, not be re-used from the source language. Re-timing captions from the original track is one of the most visible errors in a localized ad.
The mechanics of getting all of that into the file are covered in the captions post.
A Practical Localization Order of Operations
- Pick markets by unit economics, not by language count. Six markets you can support in customer service beat twenty you cannot.
- Localize the offer before the script. Price, shipping, guarantee and any regulated claim. If the offer does not work in a market, no amount of creative fixes it.
- Write the script natively per market from the same brief and the same persona, rather than translating the English one.
- Normalize numbers, units and acronyms into spoken form per language, before generation.
- Cast a creator and a voice per market.
- Generate, caption from the localized audio, and check the first three seconds in each language - the hook is the only part where a localization error is fatal rather than merely noticeable.
- Test hooks within a market before comparing across markets. Cross-market performance differences are dominated by offer and audience, not creative, so comparing a German winner against a Spanish loser tells you almost nothing.
Common Mistakes
Translating the winning script rather than rewriting from the brief. You inherit English sentence structure, English idioms and English pacing, which is what makes translated ads sound translated.
Fix: keep the brief, persona and awareness stage constant, and write the script fresh in each language.
Reusing the source language's word timings for localized captions. The captions drift against the audio within a few seconds and the ad reads as broken.
Fix: transcribe the localized audio and derive word timings from that track.
One caption font for every market. Missing glyphs are burned in permanently and cannot be fixed downstream.
Fix: verify glyph coverage per language before rendering, and fall back to a font that has the script rather than to a default that does not.
Judging a market off one localized ad. A single creative in an unfamiliar market conflates creative quality, offer fit and audience targeting into one number.
Fix: launch at least three genuinely different creatives per market so within-market variance is visible before you draw a conclusion about the market itself.
FAQs
Is dubbing ever the right call?
Yes - when the on-screen person is the point. A founder, a named expert or a known spokesperson cannot be recast, so dubbing or subtitling is the only option, and subtitling is usually the more honest of the two. For an anonymous UGC creator, though, there is nothing to preserve: recasting per market costs nothing and buys everything.
How many languages is realistic to run at once?
The generation is not the constraint. Support, returns, payment methods and legal review are. Most teams over-extend on languages and under-extend on creative variety within each one, when the reverse produces better results.
Do I need a separate creator per language, or per region?
Per market, which is often narrower than per language. Spanish for Spain and Spanish for Mexico differ in accent, vocabulary and price framing enough to be worth splitting once volume justifies it. Start at language level and split the markets that earn it.
What about regional accents within one language?
They matter more for credibility than for comprehension. A neutral accent is a safe default; a regional one can outperform it when the audience is regional, and can underperform badly when it is not. Treat it as a testable variable rather than a default.
How does siriusly.ai handle this?
Each language is a separate native render, not a dub: the script is written for that language from the shared brief, numbers and acronyms are normalized into spoken form before generation, the creator and voice are cast per market, and captions are timed from that language's own audio. The reel length is allowed to differ per language rather than being forced to match the source.