Technical

Audio Description Control Track Explained

What is an audio description control track?

An audio description control track is a two-channel signal in which channel one carries the mono description narration and channel two carries fade and pan metadata as a modulated audio signal. A receiver or playout system decodes channel two and ducks the programme audio automatically while each description plays.

The reason this exists is worth stating plainly, because it is the thing that makes the format make sense: the mix has not happened yet. The description and the instructions for how to blend it are being shipped together so that something downstream can perform the mix at the last possible moment.

That is a deliberate design choice with real consequences, and it is the first thing to establish on any audio description project — before scripting, not after.

What Is Actually in the Two Channels

  • Channel 1 — the mono audio description narration, and nothing else. No programme audio, no music.
  • Channel 2 — fade and pan data, encoded as a modulated audio signal rather than as metadata in a sidecar file.

Encoding the control data as audio is the unusual part, and it is intentional. Audio survives the broadcast chain. A pair of audio channels passes through routers, encoders, MXF wrappers, and transmission systems that would strip, ignore, or reorder a metadata track. The instructions travel in the one container the infrastructure is guaranteed not to discard.

The decoder reads channel two and applies the fade values to the main programme audio and the pan values to the narration. Ducking is therefore deterministic: it happens at the same instant, by the same amount, on every playback. This is categorically different from a dynamics processor listening for speech and reacting, which is what you get if someone attempts to recreate the effect without the control data.

Broadcast Mix vs. Receiver Mix

A control track only matters in one of these two models. Knowing which one your delivery spec is written against answers most other questions.

Broadcast mix Receiver mix
What you deliver One finished track, description already mixed in Narration plus control data; programme audio stays separate
Who mixes You, in post The viewer's receiver or set-top box
Control track needed No Yes
Bandwidth A full additional audio programme One mono channel plus control data
Viewer control None — fixed balance Description level can be adjusted
Risk None after delivery Depends on correct receiver implementation

The trade-off is control against flexibility. Broadcast mix guarantees exactly what the audience hears, at the cost of bandwidth and of giving viewers no ability to adjust the balance. Receiver mix is efficient and accessible — a viewer who finds the description too quiet can raise it — but moves part of the outcome into equipment you do not control.

The Standards Involved

Reference What it defines
BBC R&D WHP 198 The audio description studio signal — how fade and pan data is encoded on the second channel
ETSI EN 300 468 DVB service information, including the component types that distinguish receiver-mix from broadcast-mix audio description
ETSI TS 101 154 DVB audio and video coding, including receiver-mix supplementary audio
ESEF The European exchange format commonly used to hand description between facilities

Two practical notes. The DVB signalling and the control track are separate concerns — the signalling tells the receiver that a receiver-mix description stream exists, the control track tells it what to do with it, and both have to be right. And if your spec names ESEF rather than a control track, it is describing the interchange format rather than the mix model; confirm which mix mode is expected underneath it.

Do You Actually Need a Control Track?

Often not. It is specific to receiver-mix delivery, and much of the market is not receiver-mix.

Destination Typical model Control track?
US broadcast Pre-mixed secondary audio programme (SAP) No
European broadcast Frequently receiver-mix Yes
Streaming platforms Fully mixed secondary audio track No
VOD via a broadcaster Varies by platform Check the spec
Web and e-learning Mixed file, or extended description No

For what the law requires rather than what the file looks like, see our audio description compliance guide. In the US the FCC quota is 87.5 hours of described programming per calendar quarter for covered broadcasters and large subscription systems, and the phase-in reached Nielsen DMAs 111–120 on 1 January 2026 — which is why stations that never previously commissioned description are now asking these questions.

How to Produce a Control Track

  1. Confirm the mix model first. Receiver-mix or broadcast-mix, and which reference the spec names. Everything downstream depends on this, and discovering it late means re-exporting.
  2. Script the description against timecode. Descriptions have to fit the dialogue gaps. Where no gap is long enough, the choice is to cut the description or move to extended audio description — which is not available for linear broadcast.
  3. Voice it with a narrator or a synthetic voice. Closed Caption Creator offers 100+ voices, and our ElevenLabs integration for higher-end synthesis.
  4. Export narration plus timed fade and pan metadata. This is the step that distinguishes a control track workflow from ordinary description: you are exporting the mix instructions, not a mix.
  5. Encode the WHP 198 signal in a tool that produces it — Engine Desktop below — yielding a broadcast WAV with narration on channel one and control data on channel two.
  6. Verify by decoding. Play the encoded file through a decoder or receiver and listen to the ducking. Do not assume the encode worked because the file exists.

Closed Caption Creator handles steps two through four inside the same project as the captions, on the same timeline, through the Audio Description Plugin. Keeping caption events visible while writing descriptions makes gap-finding substantially faster than working from a waveform alone.

Encoding with Engine Desktop

Engine Desktop by Emotion Systems showing an audio processing workflow

Engine Desktop by Emotion Systems

Emotion Systems' Engine Desktop is a file-based audio processing tool used in broadcast and post facilities, and its support for the BBC audio description studio signal is what makes it the practical encoding step here. Import the export from Closed Caption Creator and it produces the encoded broadcast WAV.

What makes it suitable for this specifically:

  • Workflows built in a graphical interface rather than scripted
  • Manual and batch processing — the same workflow handles one file or a quarter's programming
  • Windows, macOS, and Linux
  • Faster-than-real-time processing
  • Broad media format support, so it fits into existing plant workflows

The division of labour is worth noting when planning a toolchain: the description editor owns the creative and timing decisions, the audio processor owns the encode. Tools that attempt both tend to be weaker at one.

When the Description Does Not Duck

The programme audio stays at full level

Usually channel two was lost or replaced. A chain that resamples, transcodes, or normalises audio can damage the modulated control signal while leaving it apparently present. Check that the delivered file still carries the control data as encoded, not merely that it has two channels.

Channels arrive swapped

Narration on channel two and control data on channel one produces no ducking and an audible tone where the description should be. Channel order is part of the spec; confirm it rather than inferring it from a waveform display.

Ducking happens at the wrong moments

Almost always a frame rate or timecode mismatch between the description export and the programme master — typically 29.97 drop frame against non-drop, which drifts progressively rather than failing outright. See drop frame timecode explained.

It works in the facility but not on air

Suspect the DVB signalling rather than the audio. If the stream is not flagged as carrying receiver-mix description, a correct control track will never be decoded. This is a playout configuration issue, not a post one.

Frequently Asked Questions

What is an audio description control track?

An audio description control track is a two-channel signal in which channel one carries the mono description narration and channel two carries fade and pan metadata as a modulated audio signal. A receiver or playout system decodes channel two and ducks the programme audio automatically while each description plays.

What is the difference between broadcast-mix and receiver-mix audio description?

Broadcast-mix delivers a single finished audio track with the description already mixed against the ducked programme audio. Receiver-mix delivers the narration separately with control data, and the viewer's receiver performs the mix. Receiver-mix needs less bandwidth and lets viewers adjust the balance; broadcast-mix guarantees what the audience hears.

What is BBC WHP 198?

WHP 198 is the BBC R&D white paper defining the audio description studio signal: how fade and pan instructions are encoded as a modulated audio signal on a second channel alongside the narration. It is the reference most European audio description control track workflows implement.

Do I need a control track to deliver audio description?

Only if the delivery spec asks for receiver-mix. US broadcast typically takes a pre-mixed secondary audio programme and streaming platforms usually want a fully mixed secondary track, neither of which needs a control track. European receiver-mix workflows and some VOD platforms do require one.

How do you create an audio description control track?

Script and voice the description in an audio description editor, export the narration together with its timed fade and pan metadata, then process that export in a tool that encodes the WHP 198 signal. Closed Caption Creator handles the scripting and export; Emotion Systems Engine Desktop produces the encoded broadcast WAV.

How does the receiver know when to duck the programme audio?

From the control data itself, not from detecting speech. The encoded fade and pan values tell the decoder exactly when to attenuate the main programme, by how much, and where to place the description in the stereo field, so ducking is deterministic and identical on every playback.


Delivering described programming? See the Audio Description Plugin or talk to our team.


Resources

Blog Article

The Best Audio Description Software for Broadcast

Read Now

Blog Article

Audio Description Requirements: 2026 Compliance Guide

Read Now

Solution

Audio Description Plugin

Learn More

Blog Article

Extended Audio Description

Read Now

Video Course

Audio Description for Beginners

Watch Now

Solution

Descriptive Video Transcripts

Learn More
Closed Caption Creator

Try it free for 7 days

Create closed captions, subtitles, transcripts, and audio descriptions in one application. No credit card required.