16 terms
Localization glossary
Every term you'll meet when dubbing, translating, or subtitling content with AI — defined in plain language.
- AI dubbing
- Replacing a video's spoken audio with speech in another language generated by artificial intelligence. The pipeline transcribes the original speech, translates it, synthesizes the translated lines as audio, and merges the result back into the video. AI Dubbing tool →
- Voice cloning
- Recreating a specific person's voice with AI. Dublyx samples about 30 seconds of the original speaker and reuses that voice for the translated speech, preserving tone, pace, and emotion.
- Lip-sync (AI)
- Re-synthesizing a speaker's mouth movements frame-by-frame so they match the dubbed audio. Without lip-sync, dubbed video keeps the original mouth movements, which visibly mismatch the new language.
- Transcription
- Converting speech in an audio or video file into written text using automatic speech recognition (ASR). Transcription is the first step of every dubbing and translation job.
- Source language
- The language spoken in the original file. Dublyx can detect it automatically, or you can set it manually for recordings that mix several languages.
- Target language
- The language you want the output in. Dublyx supports 20+ target languages, and AI Dubbing can produce up to 4 target languages from a single upload. Supported languages →
- SRT (SubRip Subtitle)
- The most widely supported subtitle file format: numbered blocks of text with start and end timestamps. Players, YouTube, and editing tools load SRT files alongside the video, and viewers can toggle them on or off.
- VTT (WebVTT)
- A subtitle format designed for the web, similar to SRT but with support for styling and positioning cues. Used by HTML5 video players.
- Burned-in subtitles
- Captions rendered permanently into the video frames (also called hardcoded or open captions). They display on every player and social feed, but can't be turned off or edited without re-rendering the video. Video Translation tool →
- CPS (characters per second)
- A subtitle readability measure: how many characters a viewer must read per second a subtitle is on screen. High CPS lines are flagged during translation so they can be shortened.
- Text-to-speech (TTS)
- Generating spoken audio from written text with AI voices. Dublyx ships 18 professional Kokoro TTS voices across American, British, French, Italian, Japanese, and Chinese accents.
- Speaker diarization
- Detecting who speaks when in a recording. It lets multi-speaker videos be dubbed with a separate cloned voice per speaker, so conversations stay natural.
- OCR (optical character recognition)
- Reading text out of images. Manga translation uses multi-engine OCR to extract dialogue from speech bubbles in Japanese, Korean, and Chinese comic pages.
- Typesetting (manga)
- Placing translated text back into a comic page's speech bubbles with an appropriate comic font, size, and alignment — the final step of manga translation. Manga Translation tool →
- Onomatopoeia (SFX)
- Sound-effect words drawn into manga pages (crash, whoosh, doki-doki). They're detected separately from dialogue and translated to culturally equivalent effects — or kept as originals.
- Custom glossary
- A user-defined list of term translations (brand names, technical jargon, character names) that the AI must respect during translation, ensuring consistent wording across every job.
See these concepts in action:
How Dublyx works