The Polyglot journal
How to caption sound effects, music and speakers
A speech transcript can miss the bell that prompts a reply or the voice outside the frame. Add the audible information a viewer needs, with evidence for every label.

Caption what the soundtrack contributes
Caption a sound when it helps a viewer understand the action, a reaction, the atmosphere or who is speaking. Describe what is audible, identify the source only when it is established, and synchronize the words with the event. A transcript containing all the dialogue can still leave those connections missing.
W3C’s captions guidance includes both speech and non-speech audio information needed to understand the content. This guide focuses on that editorial review: deciding what the words should convey. Work from a final recording you own or have permission to caption and distribute, including permission for its music and any translated version. Keep the original media and a separate working copy of the timed text.
Listen once for events, not just words
Before editing individual cues, listen through the finished video and make a short event log. Note the time, the audible event, why it matters and anything that needs confirmation. For a fictional clock-repair film, a bell from another room might explain why the presenter stops speaking. A barely audible chair movement might add nothing. The choice depends on the scene, not simply on whether a sound exists.
Use the log to separate evidence from an interpretation. Hearing a bell does not establish that a customer has arrived. Hearing a low rumble does not establish an approaching storm. If the source is uncertain, describe the sound at the level you can verify and keep a question in the review notes. Do not publish a more elaborate explanation because it fits the story.
Fictional clock-repair film
00:08 Bell rings offscreen
Presenter pauses
and turns
Source not yet shown
Describe the bell;
do not invent a visitorChoose a short, specific sound description
DCMP’s Sound Effects and Music guidance recommends descriptions that help the audience understand or enjoy the material, using specific wording where possible. Apply that principle by naming the audible event rather than writing a vague label such as a noise. In the clock film, [bell rings] says more while staying within the evidence. If the picture and soundtrack establish a particular clock, a more specific description may be appropriate.
For a sustained event, wording such as [mechanism rattling] can communicate continuation. For a single event, [bell rings] describes the change. Match the caption to the actual sound; do not extend it across a silent passage just to fill space. Repeated labels should earn their place by conveying a meaningful return or change, rather than making every moment compete with the dialogue.
Keep descriptions distinct from spoken words using the convention required by the recipient. Square brackets are a common sound-description convention and are used in our examples. A broadcaster, education service or streaming destination may prescribe its own punctuation, placement and timing rules. Follow that delivery guide instead of treating an example here as a universal specification.
Identify a speaker without revealing more than the film
Watch the speaker changes with captions visible. If the picture makes the speaker clear, a repeated name can be unnecessary. When a reply comes from outside the frame or a cut makes attribution ambiguous, a label can preserve the connection. DCMP’s speaker-identification guidance says not to reveal a name before the audio or an onscreen introduction establishes it.
In our fictional film, Mara has already been introduced. When she answers from the next room, a label using her established name is helpful. If the voice has not been identified, use a neutral, consistent label appropriate to the delivery guide, such as a numbered speaker. Do not guess age, gender, a job or a personal identity from the sound of a voice.
Keep a small reviewer note mapping each label to the evidence that establishes it. The name should stay consistent across cuts and translated tracks. If two people overlap and the words cannot be recovered confidently, return to the audio or ask the producer; assigning a confident label does not resolve uncertain speech.
Describe music without adding an opinion
Music can establish an arrival, a change in pace or tension that the dialogue never names. Use a concise description of the audible character when it contributes to the scene. In the clock film, [slow music-box melody] would be useful only if that description is supported by what is heard or shown. Do not name an instrument, recording or composer merely because it sounds familiar.
DCMP advises objective music descriptions. Choose audible qualities over praise such as beautiful or wonderful. Check any requested song title, performer credit or lyrics against authorized material and the recipient’s captioning guide. Our example contains no lyrics. A music label should not displace intelligible dialogue or cover an important visual detail; review both together at the destination’s actual display size.
Keep the cue example and the evidence together
This original SRT example shows a bell followed by an identified offscreen reply. Mara was introduced earlier, the bell is heard at eight seconds, and her line starts at eleven seconds. The timing and dialogue are invented solely to explain the format; they are not a tested Polyglot transcription or a template to paste into another recording.
The sound cue has its own interval because no speech occurs during that event. In your footage, check the actual start, duration and relationship to neighboring dialogue. If a sound and speech happen together, use the destination’s conventions and review the combined reading load. Do not move an event into an unrelated speech cue simply because that cue is convenient to edit.
1
00:00:08,000 --> 00:00:10,000
[bell rings]
2
00:00:11,000 --> 00:00:14,000
(Mara)
I heard it from the next room.Use Polyglot for the parts it supports
Polyglot can create a timed speech draft from authorized media, or import an existing SRT or WebVTT file. You can review and edit cue text, translate it and export a separate subtitle file. Speech recognition is a starting point: Polyglot does not promise automatic speaker identification or complete sound-effect and music descriptions.
If an existing cue already covers the correct event, review its wording and any necessary label against playback. If the file lacks a sound-only interval, use a caption editor that can add and time individual cues, then import the reviewed SRT or WebVTT. Do not stretch nearby dialogue to stand in for missing cue-authoring work. Polyglot’s uniform timing shift moves the file together; it does not create a new interval for a bell.
After changing source wording, review affected translations as well. Check that speaker labels stay attached to the right voice and that translated sound descriptions still express the same event. The exported file remains separate from the video; loading it into a player, publishing a caption track or making a burned-in copy is a separate delivery step.
Review with sound, then without it
First watch with sound to check every description against the audible event and every label against the established speaker. Resolve uncertain notes before approval. Then watch with sound muted to find missing connections: a reply with no apparent trigger, a sudden reaction with no explanation, or a speaker change that the captions leave ambiguous. A muted pass reveals gaps; it cannot prove that the words match the audio.
Finally load the exported file into the intended player. Check that brackets and names display correctly, descriptions have enough reading time, and captions do not hide essential picture information. Follow through on meaningful events, rather than treating a working caption toggle as proof of completeness. Important information conveyed only by the picture may need a separate description approach; adding sound captions alone does not make every part of a video accessible.
Sources & further reading
W3C WAI: Captions/Subtitles — speech, non-speech audio and review
DCMP Captioning Key: Sound Effects and Music
DCMP Captioning Key: Speaker Identification
Published by Polyglot using an AI-assisted editorial workflow. How we prepare and update our guides.
Keep exploring.
Try it on your Mac.
Find the right tools for your words and video, with processing on your Mac.
Download nowApple Silicon · macOS 14.4 or later