StemVocalPricing

Podcast transcription

Turn an episode into a transcript you can check against the audio, then publish it with the episode or in your show notes.

Transcribe a recording

Transcribe an episode

Start from the finished episode, after editing, so the transcript matches what listeners hear. Export it as MP3, M4A or WAV, open the transcription studio and choose the file. Pick the spoken language, or let StemVocal detect it.

An episode up to 30 minutes can use one of your 2 free jobs for the day. Longer episodes, up to 3 hours, use one credit for each started 30 minutes, so a 75-minute episode uses three. Credit packs, the Transcription plan and Unlimited are on the pricing page.

For a show with hosts and guests, choose how many people speak before you start. The number is a maximum, up to 8. Speaker labels use credits, and they are estimates: they work best when people take clear turns and need a careful check when voices overlap or sound alike. Rename each label to the person's name once the transcript is ready.

Check names, numbers and terms

Podcasts are full of guest names, brands and in-jokes that automatic transcription rarely gets right the first time. Click any word to hear that moment, slow the playback down for a fast talker, and search the transcript for a phrase. When a name is misheard the same way all the way through, open Find & replace, check the match count and fix every instance at once. One Undo restores the whole replacement.

When people talk over each other, the transcript usually keeps one voice. Listen again around interruptions and laughter, and around ads or music beds, before you publish.

Choose the file to publish

Where the transcript goes Download What to know
Show notes or an episode page Word or plain text Turn on timestamps to point readers to a moment
A podcast app that reads transcript files VTT or SRT Words and cue times; speaker names are not included
The video version of the episode SRT or VTT Import it as a caption track and check the timing
A spreadsheet of quotes and timings CSV Start and end seconds, speaker and text for each turn

Word, text and print files can start each turn with a timestamp, written as [hh:mm:ss] Speaker: text. Subtitle cues are timed from the words themselves: each lasts at most 6 seconds, holds at most 80 characters and never overlaps the next one.

Publish it with your episode

Apple Podcasts creates transcripts automatically after a new episode is published. Listeners see them on iOS 17.4 or later, for podcasts in English, Danish, Dutch, Finnish, French, German, Italian, Norwegian, Portuguese, Spanish and Swedish. To show your own transcript instead, change the transcript setting in Apple Podcasts Connect (the Availability tab of your show, or a custom setting for one episode) to display transcripts you provide. Apple then reads the file from your RSS feed's transcript tag; for subscriber episodes, you upload the VTT or SRT file in Apple Podcasts Connect. Apple labels these transcripts "Provided by" your show, and files that do not meet its quality standards are not displayed.

Apple notes that a VTT file can carry each line's speaker name. StemVocal's VTT and SRT files contain the words and timing without names, so speakers are not marked line by line in Apple's transcript view. Whichever transcript you use, Apple recommends putting your hosts' and guests' names in the show and episode descriptions so that its automatic transcripts spell them correctly.

Other apps and hosts use the Podcasting 2.0 podcast:transcript tag in the feed. It needs two attributes: url, the address of the transcript file, and type, its format: text/vtt for VTT, application/x-subrip for SRT, or text/plain for plain text. The optional language attribute names the transcript's language, and rel="captions" marks a timed captions file. A feed can list several transcript files, one per format. Most people add transcripts through their hosting service, so check whether yours supports them before editing the feed.

Keep your transcripts

Finished transcripts save in this browser under Saved transcripts, so you can reopen an episode and correct it later without another job. Uploaded recordings stay available for 24 hours and are normally deleted within the following 24; you can delete one sooner from the editor. Download the files you publish. For interviews recorded for a single episode, see interview transcription.

Sources: Apple Podcasts for Creators, "Transcripts on Apple Podcasts"; Podcast Index, Podcasting 2.0 namespace, podcast:transcript tag. Checked October 11, 2026.

Good to know

Can I transcribe a podcast for free?

Yes, for episodes up to 30 minutes. You get 2 free StemVocal jobs each UTC day without an account, shared with the other StemVocal tools. Longer episodes, up to 3 hours, use one credit per started 30 minutes, and speaker labels use credits too.

Will the transcript show who is speaking?

With speaker labels, which use credits. Choose the most voices you expect, up to 8, then rename each label to a host or guest. Labels are estimates and work best when people take clear turns. Text, Word, print and CSV downloads include the names; SRT and VTT files hold the words and timing only.

Which transcript file does Apple Podcasts accept?

Apple Podcasts reads transcripts you provide in VTT or SRT format. It takes them from your RSS feed's transcript tag, or as an uploaded file for subscriber episodes, once the show or episode is set to display transcripts you provide.

Can I translate a podcast transcript?

A paid transcript can be translated from English into 15 languages, or from any of them into English, and downloaded as text or subtitles. It is machine translation, so have a fluent speaker review it before you publish.