The Polyglot journal
Turn SRT subtitles into a readable transcript on Mac
A subtitle export can become a useful reading copy. Preserve the reviewed words, choose time labels deliberately, and check the context that a plain text file cannot supply.

Start with the reviewed subtitle file
To convert SRT subtitles into a readable transcript on Mac, open the reviewed track in Polyglot and export it as Text transcript. Choose whether to include timestamps, then check the resulting text for paragraph breaks, speakers and missing context. Keep the timed subtitle file: a text transcript is a separate deliverable and cannot replace captions in a video player.
This workflow starts with text you already have. It does not require transcribing the recording again. Use an SRT or VTT you own or have permission to process and share, and confirm that it belongs to the final media edit. If the subtitle wording is still uncertain, resolve that first; conversion carries those words forward rather than checking them against the recording.
Choose the track and the kind of transcript
Open the saved original or translated track you intend to deliver. For an imported subtitle file, check the language and compare its opening and closing passages with the matching recording when available. A filename such as studio-demo-v4.en.srt helps identify the revision, but the name alone does not establish that it is the right cut.
Decide what the recipient needs before removing timing. A reading copy may work best as paragraphs without time labels. A review copy can retain cue-start timestamps so someone can locate a disputed phrase. A descriptive transcript also needs relevant visual information, which must be prepared separately. An interactive transcript that jumps to the video when clicked is a player feature; a TXT export does not provide that behavior.
Export text, with timestamps only when useful
In the project’s Export menu, choose Export subtitles or transcript… and select Text transcript as the format. Check the Language selection. Leave Include timestamps off for a reading copy, or enable it for a copy with cue-start times. Export the file, or use Copy text when you need to paste the result into another document.
Polyglot writes UTF-8 text. It joins the line breaks within each cue into ordinary spacing; without timestamps, it groups the cue text into paragraphs using timing gaps and length. These are starting paragraphs, not an understanding of the subject or speaker turns. With timestamps enabled, each cue gets its start time rather than the full SRT start-and-end interval.
Save the text under a new filename, such as studio-demo-v4.en.transcript.txt, beside the source subtitles. Keep an untouched export and make editorial changes in a separate reading copy. Renaming an SRT to .txt does not perform this conversion: its cue numbers and timing lines would still be part of the content.
Example: preserve the words, then add context visibly
The fictional craft demonstration below is original sample text, not an accuracy test. Its first sentence spans two cues. Removing subtitle timing should join those pieces without dropping the quantity or the exception. The third cue begins after a longer pause and belongs in a new paragraph.
The plain reading copy preserves what was said. The optional contextual version adds a verified speaker label and a bracketed visual observation. Those additions are manual editorial work, not generated by transcript export. If the actual video did not show a narrow blue strip, that description would be inappropriate; review the recording before adding it.
Source SRT
1
00:00:01,000 --> 00:00:03,000
Fold the strip
into three sections,
2
00:00:03,100 --> 00:00:05,500
but leave the last fold open.
3
00:00:08,000 --> 00:00:10,000
Now turn it over.
Plain reading copy
Fold the strip into three
sections,
but leave the last fold open.
Now turn it over.
Optional manual context
Presenter: Fold the strip
into three
sections, but leave the last
fold open.
[The presenter turns a narrow
blue
paper strip over.]
Presenter: Now turn it over.Review paragraphs, speaker turns and repetitions
Read the text from beginning to end without watching the video. Mark places where a sentence stops unexpectedly, a new speaker appears to continue someone else’s thought, or a reference such as “this one” has no clear meaning. Then return to the recording to resolve each issue. A pause is not necessarily a new topic, and a quick exchange may contain several speakers without a long pause.
Combine fragments that form one sentence and separate real changes of speaker or topic. Keep a verified name or a consistent neutral label where needed. Do not infer identities from the voice or invent a name because a paragraph needs a heading. Polyglot does not automatically identify speakers for this workflow.
Repeated words need judgment. A speaker may deliberately repeat an instruction; two people may say the same thing; or the source captions may contain accidental duplication. Compare the source cues and audio before removing anything. Likewise, keep meaningful sound information already present in reviewed captions. W3C’s caption guidance treats relevant non-speech audio as part of the information viewers may need.
Keep descriptive additions separate from speech
A speech-only export can leave out the very thing the presenter is pointing at. W3C’s media guidance encourages making essential visual information available beyond the picture. For a descriptive transcript, review the actual video for necessary on-screen text, demonstrations and changes that the audio does not explain. Describe only what you can establish from the approved media.
Label an editorial addition so a reader will not mistake it for spoken dialogue. In the example, brackets distinguish the paper-strip action from the presenter’s words. Keep added descriptions factual and concise. If a label on screen is unreadable, record that uncertainty for the content owner instead of guessing its wording.
Keep these additions in the transcript document unless you are deliberately revising the caption deliverable too. Pasting a longer descriptive paragraph back into a timed cue can make the subtitle impossible to read in its original interval. Export does not generate image descriptions or assess whether a transcript meets an accessibility requirement; that needs a review of the actual media and intended use.
Check the copy where it will be read
Open the exported TXT in the editor or publishing workflow the recipient actually uses. Check accented names, apostrophes, paragraph separation and optional time labels. If you paste into a website or document editor, check the pasted result too: the destination may apply its own spacing and styles. Keep the language explicit, particularly when the project contains both original text and a translation.
If you need a polished web transcript, add a clear title and helpful headings in the publishing system. W3C recommends making transcripts easy to find from their media; a visible link beside the video is more useful than an unexplained download elsewhere. A plain-text export supplies content for that work, not a finished HTML page or a linked media player.
Finally, compare the reading copy with the reviewed source track. Check every passage, paying particular attention to numbers, negatives, speaker changes and the opening and closing lines. Confirm that no cue text disappeared during cleanup, that deliberate additions are identifiable, and that unresolved wording is held for review. If the video or source subtitles change later, revisit the transcript rather than assuming the earlier copy is still current.
Hand over a small, traceable set of files
Deliver the final reading copy with its language, media revision and review status. Retain the original timed track and the untouched text export so another reviewer can distinguish conversion from later editing. For example, keep studio-demo-v4.en.srt, studio-demo-v4.en.transcript.txt and a separately named reviewed document together.
Before sharing, confirm that the recipient has the intended version and that the sharing permission covers the text as well as the recording. Keep private review notes out of the public transcript. The useful result is a readable account whose words and additions can be checked against a known source, while the original subtitle timing remains available for video delivery.
Sources & further reading
W3C WAI: creating and publishing transcripts from captions
W3C WAI: making visual information available in media
W3C WAI: captions include meaningful audio information
Published by Polyglot using an AI-assisted editorial workflow. How we prepare and update our guides.
Keep exploring.
Try it on your Mac.
Find the right tools for your words and video, with processing on your Mac.
Download nowApple Silicon · macOS 14.4 or later