Caption the actual soundtrack and make speaker attribution clear where the audience needs it. Two simultaneous voices should not become one unattributed sentence. Review the overlap against the audio, use consistent speaker labels when appropriate and test whether the rendered cues remain readable. Do not infer the speaker solely from who is nearest the camera or whose replacement portrait is most prominent.
Map the audio before writing labels
Review the clean dialogue tracks if available, then listen to the final mix. Note who speaks, where speech overlaps and which words are actually intelligible. A proposed script may differ from the performed or generated soundtrack. Captioning cannot repair a scene in which the wrong character audibly says the line; flag that as a separate edit problem.
W3C WAI's caption guidance explains that captions convey relevant audio information as well as speech. Speaker identification should support understanding without adding unsupported detail. If a narrative intentionally hides a voice's identity, use a suitable neutral attribution rather than reveal it early.
Test an overlap in context
- Mark the start and end of the overlapping speech.
- Transcribe the audible words and verify the speaker assignment.
- Choose consistent labels and a readable cue arrangement.
- Review the actual rendered overlap with sound off.
- Listen again with sound to check timing and meaning.
For a proposal film, an excited interruption may carry emotional meaning. Do not silently convert it into a calm sequential exchange simply to make the subtitles tidy. If words are unclear, obtain an approved correction or use the appropriate uncertainty convention for the destination.
Keep labels stable across scenes
Maintain the same approved speaker name or descriptive label throughout the film. A label that changes from a person's name to a costume color can confuse viewers after a character replacement or wardrobe change. Review names against the final credits and the user's approved wording.
Should every cue repeat the speaker's name?
Not necessarily. Use identification where needed to follow the dialogue, then assess readability in context. Avoid unnecessary clutter that leaves less time to read the spoken words.
Can automatic transcription settle the attribution?
Treat it as a draft. Review the actual voices and overlaps. A plausible transcript with swapped names can change the scene's meaning even when every word is spelled correctly.
Sources and scope
This is an original Imagild Editorial workflow guide, assisted by AI and reviewed against the linked sources on October 12, 2026. External tool procedures do not imply that Imagild provides those tools or guarantees their results.