The Polyglot journal
How to write a descriptive transcript for a video
Accurate speech is only part of a video. Build a descriptive transcript that explains important actions, labels and visual changes without turning observation into invented dialogue.

Give the reader the information the picture carries
To write a descriptive transcript, start with reviewed speech and meaningful sound information, watch the final video for essential details that the audio leaves unexplained, and add clear descriptions at the relevant points in the text. Then have someone read the document without relying on the video. The reader should be able to follow the same important actions, references and changes.
W3C WAI distinguishes a basic transcript of audio information from a descriptive transcript that also conveys necessary visual information. A presenter saying “use this row” may be transcribed perfectly while leaving a reader unable to identify the row. That missing reference is the problem this guide addresses.
Use this workflow for recordings you own or have permission to process and publish. A descriptive transcript is a separate reading document. It does not put captions on a video or create a spoken audio-description track, and it should not be treated as proof that every accessibility requirement has been met.
Keep a stable source and a separate working document
Identify the exact final video revision and language. Start with reviewed speech, speaker labels where needed, and meaningful non-speech sounds. Resolve uncertain words before using them as the foundation for further editing. If you already have accurate captions, the existing subtitle-to-transcript guide explains the conversion step.
Polyglot can export the words from a subtitle track as a text transcript, with optional cue-start timestamps. That export supplies a starting document; it does not inspect the picture or generate visual descriptions. Keep the original subtitle file and untouched text export alongside a separately named descriptive draft.
Work in a document editor or publishing system for the additions. Do not paste long visual descriptions back into the timed subtitles by default: those cues were prepared for a different reading situation. Retaining the source files also lets a reviewer distinguish the recorded speech from later editorial work.
Make a visual-information inventory before writing
Watch the video in short passages and note information that is visible but not adequately conveyed in the audio. Useful categories include written labels and quantities; actions and their results; relationships in a diagram; and context needed to understand who or what the speaker refers to. Include silent passages rather than looking only where there are subtitle cues.
Give each observation a decision. Add it when it is necessary to understand the passage. Mark it already spoken when the audio explains it sufficiently. Omit it when it has no bearing on the content. Mark it unresolved when the picture is too unclear to support a trustworthy description. These are editorial decisions, not classifications supplied by Polyglot.
Avoid inventorying every visible object. In a planting demonstration, which tray row receives the seeds may matter; the decorative mug behind the presenter probably does not. In another video that same mug could be the subject. Relevance follows the task and the recording, not a fixed list of objects to describe.
For charts, preserve the relationship that matters, together with labels, units or values needed to understand it. Do not infer a cause from a line going up. For a demonstration, describe the action and visible result instead of an intention you cannot observe.
Worked example: resolve “this row” without guessing
The following seed-tray lesson is a fictional example written for this guide. Imagine that the approved video shows a two-row tray: the near row is labeled A and the far row B. The presenter points to row B while speaking. Later, the camera shows one seed in each cell of that row. A small variety name is too blurred to read.
The review note links each missing detail to its place in the video. The variety name stays unresolved; it must be clarified with the content owner and checked against an approved source before it can be included. If that name is essential to the lesson, hold that part of the handoff rather than guessing.
Visual review notes
00:06 — Row labels A and B
Decision: Add the layout
00:09 — Points to far row B
Words: “Use this row.”
Decision: Add the reference
00:14 — One seed per cell
Decision: Add visible result
00:20 — Blurred variety name
Decision: Resolve with owner
Background mug
Decision: Not neededPut each addition where it becomes meaningful
Introduce the tray layout before the reader needs to understand the pointing gesture. Put the visible result where it occurs, after the instruction. Do not move a later result earlier simply because it makes a neater paragraph: that can change the sequence of a demonstration or reveal an outcome before the video does.
Keep the spoken sentence intact, and distinguish the added explanation using consistent labels or brackets. In the example below, the bracketed lines are manual descriptions, not words spoken by the presenter. “Presenter” is a neutral label; it does not invent a personal identity.
Use precise, ordinary wording. A label copied from the image needs a spelling check; a measurement needs its unit; a changing screen state needs its before-and-after relationship. If the audio already says the necessary detail, avoid immediately repeating it. The aim is a coherent account, not alternating duplicate versions of every sentence.
[Visual: A tray has two rows.
The near row is labeled A;
the far row is labeled B.]
Presenter: Use this row.
[Visual: The presenter points
to row B, the far row.]
Presenter: One seed in each
cell.
[Visual: Each cell of row B
contains one seed.]Review once as a reader, then against the video
First, ask a reviewer to read the draft without watching. Have them explain which row is used, what action happens and what result is shown. These questions expose missing references better than asking only whether the transcript reads smoothly. In a real project, choose questions about the actual essential content.
Then compare the document with the complete final video. Check that speech and meaningful sounds remain present, that descriptions match what can be established, and that their order follows the recording. Look for unsupported motives, missing labels, accidental duplication and an unexplained “here” or “this.” Keep uncertainty visible in private review notes until resolved.
If a translated descriptive transcript is needed, review the added descriptions as well as the spoken material. Translating the subtitle track alone does not translate descriptions added later in a separate document. Check direction words, diagram labels, names and measurements with a qualified language reviewer for the subject.
This two-pass method is a proposed editorial check, not a claim that this article or Polyglot has tested your media. For consequential instructional content, involve the responsible content owner and appropriate accessibility reviewers.
Publish a readable document people can find
Give the finished transcript a clear title, language and useful section headings. W3C recommends making transcripts easy to find from their media; place a descriptive link beside the video or include the transcript on the same page. A file hidden in a downloads directory is harder to discover.
Check the actual published reading order, especially if your publishing system uses separate audio and visual columns. A narrow screen should not require the reader to jump unpredictably between disconnected passages. For a straightforward lesson, interleaved paragraphs with clearly marked descriptions are often a practical starting format.
Keep optional timestamps when they help a reader locate a passage, but do not describe plain text time labels as clickable video navigation. Interactive transcript behavior depends on the receiving player and publishing implementation. A TXT export does not provide it.
Approve the document against a named video revision
Before handoff, record the video revision, transcript language, reviewer and any unresolved essential information. Keep internal questions out of the public document. Confirm that permission to distribute covers the transcript and any reproduced on-screen material as well as the recording.
For the fictional lesson, approval means the reader can identify row B, follow the planting sequence and understand the visible result, while the essential variety-name question has been resolved or the affected delivery remains pending. A correctly exported text file alone cannot establish those things.
When the video changes, recheck the visual inventory as well as the speech. A replaced diagram or reordered demonstration can invalidate a descriptive transcript even if the spoken words are unchanged. Retain the source track, the separate descriptive document and a short revision note so the next reviewer knows what was approved.
Sources & further reading
W3C WAI: basic and descriptive transcripts
W3C WAI: describing visual information
W3C WAI: planning captions, description and transcripts
Published by Polyglot using an AI-assisted editorial workflow. How we prepare and update our guides.
Keep exploring.
Try it on your Mac.
Find the right tools for your words and video, with processing on your Mac.
Download nowApple Silicon · macOS 14.4 or later