The Best Audio Description Software for Broadcast in 2026
What do you need in audio description software?
Audio description software has to do four things: script descriptions against timecode, fit them into the dialogue gaps, voice them by narrator or synthetic voice, and output the mix your delivery spec asks for. Tools that only do the first two leave you finishing the job in a DAW.
Audio description demand is rising on a fixed schedule. The FCC's phase-in reached Nielsen DMAs 111 to 120 on 1 January 2026 and adds ten more markets every January through 2035, so stations that had never commissioned description are now doing so on a quarterly quota.
That has changed what teams need from the software. The question is no longer whether a tool can write a script — all of these can — but whether it can carry the work all the way to a deliverable file without a hand-off to a separate audio department.
Here is how six current options compare. For what the law actually requires, see our audio description compliance guide; this article is about the tooling.
Audio Description Tools at a Glance
| Tool | Platform | Synthetic voice | Human recording | Mix / encode | Best for |
|---|---|---|---|---|---|
| Closed Caption Creator | Windows, macOS, Linux, browser | Yes — 100+ voices | Yes | Mixed and discrete audio export | Teams already captioning who need description too |
| Frazier | Desktop and browser | Yes — Google, Microsoft, AWS, ElevenLabs | Yes | Script and audio export | Description specialists who script all day |
| Stellar (YellaUmbrella) | Browser | Yes | Yes | End-to-end: script, record, QC, mix | Facilities wanting one system with review built in |
| Starfish Technologies | Windows / appliance | Limited | Yes | ESEF, linear BWAV, frame rate conversion | Broadcast plants delivering ESEF |
| ADAutor | Windows | Bring your own voice packs | Yes | Script and audio export | Low-budget in-house description |
| DAW (Pro Tools, Audition, Reaper) | Windows, macOS | No | Yes | Full mixing control | Finishing, not scripting |
Disclosure: Closed Caption Creator is our own product. We have tried to describe where it is and is not the right choice.
The Four Stages Every Tool Has to Cover
Comparing audio description tools is easier once you separate the workflow into stages, because most products are strong at some and absent at others.
1. Timed scripting
Descriptions have to land in the gaps between dialogue, which means the scripting view needs a waveform, the programme video, and ideally the caption track so you can see where speech already occupies the timeline. Tools that let you write against timecode rather than a word processor save the most time here.
2. Fitting to the gap
A description that reads well but runs 1.5 seconds longer than the available gap is not usable. Good tools show the available duration and warn when the synthesised or spoken length overruns it. When no gap is long enough, you either cut the description or move to extended audio description, which pauses the video — an option that is not available for linear broadcast.
3. Voicing
Either a narrator in a booth or a text-to-speech engine. Synthetic voice quality crossed the acceptability threshold for factual and library content some years ago and is now standard for high-volume work. Drama and premium commissions still frequently specify a human narrator.
4. Mixing and delivery
The description has to be mixed against the programme audio with the programme ducked underneath, then exported in whatever the spec requires. This is the stage where tools differ most, and the one teams most often discover late.
The Tools
Closed Caption Creator
https://www.closedcaptioncreator.com/
Audio description runs as a plugin inside the same editor used for captions and subtitles, on the same timeline. That means the caption events are visible while you write descriptions, which makes gap-finding considerably faster than working from the waveform alone. It runs on Windows, macOS, Linux, and in the browser.
Key features:
- 100+ synthetic voices with multi-language support
- Caption events visible on the description timeline
- Human recording as well as synthesis
- Export as FLAC, MP3, or PCM WAV, mixed or as a discrete stem
- Local video import — media is not uploaded
- Supports an audio description control track for playout systems that duck automatically
Use it when you are already captioning and want description handled in the same project rather than in a second system with a second licence.
Frazier by Video to Voice
The most refined scripting interface of the group, and the tool most audio description writers reach for by preference. Frazier times descriptions in real time as you type and hooks into text-to-speech engines from Google, Microsoft, AWS, and ElevenLabs. It is tiered, with script export in the base version and multi-format subtitle export and premium voices higher up.
Key features:
- Real-time timing feedback while writing
- Multiple TTS engines including ElevenLabs
- Live collaborative editing on the same script
- Tiered licensing — the useful features sit above the base tier
Use it when scripting volume is your bottleneck and you have an audio path for the mix already.
Stellar by YellaUmbrella
https://yellaumbrella.com/audio-description/
A browser-based suite that covers scripting, recording, QC, and mixing in one system, so nothing has to be converted or handed between tools. Consumption-based pricing — per minute of video plus user accounts — suits organisations with uneven monthly volume better than flat per-seat licensing.
Key features:
- Script, record, QC, and mix without leaving the platform
- Multi-channel audio support
- Task management and review built in for team workflows
- Consumption pricing — cheap in quiet months, less so in busy ones
Use it when you want one platform end to end and your volume varies enough that per-minute pricing beats fixed seats.
Starfish Technologies
https://www.starfish.tv/audio-video-description/
The most broadcast-plant-oriented option here. Starfish produces ESEF files and linear broadcast WAV, and handles 24/25 frame rate conversion — the practical detail that matters for European delivery and for anything moving between film and PAL rates. Scripting and recording happen simultaneously.
Key features:
- Industry-standard ESEF output
- Linear broadcast WAV delivery
- 24/25 frame rate conversion built in
- Record and script in one pass
- Narration is recorded rather than synthesised
Use it when your delivery spec names ESEF, or when frame rate conversion between 24 and 25 is a routine part of your workflow.
ADAutor
https://www.audiodescription.info/software/adautor/
The budget option, and honest about it. ADAutor is a Windows scripting and recording tool that expects you to supply your own synthetic voice packages and do some setup before first use. For an organisation producing description occasionally and in-house, the low cost can outweigh the rough edges.
Key features:
- Low cost, Windows desktop
- Record with your own voice
- No bundled synthetic voices — you install voice packs separately
- Initial setup is more involved than the commercial alternatives
Use it when budget is the binding constraint and description is occasional rather than continuous.
DAWs and General Audio Tools
Pro Tools, Adobe Audition, and Reaper come up often, and they belong in the workflow — just not at the start of it. They give you total control over the mix, ducking, and loudness compliance, which the dedicated tools abstract away. What they do not give you is timed scripting against the programme, gap-length warnings, or synthetic voice.
Use it when you have a finished, timed script and audio, and need a precise mix that meets a loudness spec.
What You Actually Deliver
The most common late surprise in an audio description project is the deliverable, not the script. Broadly:
| Destination | Typical deliverable | Notes |
|---|---|---|
| US broadcast | Mixed secondary audio programme (SAP) or discrete stem | Delivered against the FCC quarterly quota; check whether the plant wants pre-mixed or mixes at playout |
| European broadcast | ESEF, or a control track plus discrete narration | The receiver mixes; frame rate conversion between 24 and 25 is common |
| Streaming platforms | Fully mixed secondary audio track | Loudness spec usually applies to the described mix as well as the main mix |
| Web and e-learning | Extended description, or a described video file | WCAG allows extended description; broadcast does not |
| Any of the above | Descriptive video transcript | A text alternative combining dialogue and description — increasingly requested alongside the audio |
Confirm which of these your client wants before scripting starts. Retiming a finished description to a different frame rate or re-mixing to a different loudness target is significantly more work than getting it right at the outset.
Which One to Pick
- Already captioning, want description in the same project: Closed Caption Creator.
- Scripting is the bottleneck, audio path already exists: Frazier.
- Want one platform end to end with review built in: Stellar.
- Delivering ESEF or converting between 24 and 25: Starfish.
- Budget-constrained and occasional: ADAutor.
- Script done, need a precise mix: your DAW.
As with captions, test against a real delivery early. Script and mix one short segment, deliver it in the exact format the spec names, and see what comes back before you commit a quarter's worth of programming to a toolchain.
Frequently Asked Questions
What software is used to create audio description?
Audio description is written in a dedicated scripting tool that times each description to the gaps in the dialogue, then voiced either by a narrator or a text-to-speech engine and mixed against the programme audio. Closed Caption Creator, Frazier, Stellar, ADAutor, and Starfish all cover that workflow; general audio editors like Pro Tools or Audition handle the mix but not the timed scripting.
Can synthetic voice be used for audio description on broadcast?
Yes. The FCC rules set out how much described programming must be provided, not how the narration is produced, and synthetic voice is now widely used for factual, educational, and library content. Many drama and premium commissions still specify a human narrator, so check the delivery spec before assuming text-to-speech is acceptable.
What file do I actually deliver for audio description?
Broadcast delivery is usually a mixed or discrete audio stem as a broadcast WAV, sometimes with a control track that tells the playout system when to duck the programme audio. European workflows often use ESEF. Streaming platforms typically want a fully mixed secondary audio track alongside the main mix.
What is the difference between standard and extended audio description?
Standard audio description fits entirely into the natural gaps in the dialogue, so the programme runs at its original length. Extended audio description pauses the video when a gap is too short for the description needed, which makes it longer than the source. Extended is common in e-learning and on the web, and is not usable for linear broadcast.
How many hours of audio description does the FCC require?
Covered broadcasters and large subscription systems must provide 87.5 hours of audio described programming per calendar quarter, of which at least 50 hours must be prime time or children's programming. The rules reached Nielsen DMAs 111 to 120 on 1 January 2026 and expand by ten more markets each year. Our compliance guide covers exemptions and the full market schedule.
Do I need a separate tool for audio description if I already have a caption editor?
Not necessarily. Some caption editors, including Closed Caption Creator, include audio description scripting and synthesis as a plugin, which keeps one timeline and one project for both deliverables. Separate tools are worth it when your description volume is high enough to justify a dedicated recording booth and mixing workflow.
Adding description to an existing captioning workflow? See the Audio Description Plugin or talk to our team.