Skip to main content

Voice Properties

Each Audio Description event exposes two voice-related properties that directly affect how the rendered audio sounds: Speed and Voice Style. Understanding how they interact allows you to fine-tune the delivery of each narration line so that it fits naturally within its timing window and matches the tone of the content being described.

Audio Dubbing events use the same Speed control. Voice Style is available on Audio Description events with supported Microsoft Azure Neural voices; it is not used for Audio Dubbing.

Speed​

Speed controls how fast the rendered voice-over plays for a given event, from 0.5 to 2. A value of 1 is the voice's natural pace. Increasing Speed makes the line finish sooner without changing pitch; lowering Speed stretches the line.

Text-to-speech always renders at provider rate 1. After that, Speed is the only timing control. It is shown for every voice, including ElevenLabs Eleven v4, Google Gemini, and Microsoft Azure DragonHD.

Speed becomes particularly important when the rendered audio is longer than the available event window. After the initial render, if a red DURATION indicator appears, it means the audio does not fit. Rather than always extending the out-time, you can raise Speed modestly so the line fits the current start and end. Conversely, if a window is unusually long relative to the narration, lowering Speed can fill the gap more naturally.

The Speed slider is located within the event card in the Event List. It applies only to the specific event where it is set. After adjusting Speed, click the force-render button on the event if you also changed text or voice; Speed itself is applied on playback and at mixdown.

Auto-set Speed​

For text-to-speech voices, each event includes an Auto-set Speed to match event duration, then render button (speedometer icon). Use this when you want Closed Caption Creator to pick a Speed that fits the rendered narration inside the event's current start and end time.

Auto-set Speed is available when the event has a supported TTS voice, valid timing, and text. It is not used for manual recordings. If the event has not been rendered yet, Closed Caption Creator renders it first at Speed 1, compares the audio duration with the event duration, then sets Speed (speeding up only, capped at 2). The operation is a single pass and is recorded in the undo history so you can revert it if the result does not fit the creative intent.

AI Tools β†’ Force Render Audio… and Force Render ALL Audio… offer the same choice as Fit to Event duration, which runs this Speed adjustment for the selected events (or the whole group).

ElevenLabs Model​

When an event uses an ElevenLabs voice, the dropdown under the voice name selects the ElevenLabs model for that event:

ModelBest forSpeed
Eleven v4The most expressive, natural delivery. The default for newly assigned ElevenLabs voices.Available (0.5–2). The model sets its own internal pacing; Speed still scales the rendered file.
Multilingual v2Stable, consistent long-form narration. Events created before model selection keep using this model.Available (0.5–2)
Flash v2.5Fast rendering at lower cost.Available (0.5–2)

Stability and Similarity are available for every model. Changing the model marks the event for re-rendering. Speed is shown for every model.

Voice Style​

Voice Style is an expressive modifier available on supported Microsoft Azure Neural voices in Audio Description events. Where a standard synthetic voice delivers text in a neutral, consistent tone, a voice style shapes how the narration is performed β€” options vary by voice but typically include styles such as cheerful, empathetic, newscast, or documentary. Not all Microsoft voices support every style, Microsoft DragonHD voices infer their delivery from the text and do not accept a style, and Amazon, Google, and Gemini voices do not currently support styles.

Audio Dubbing does not show a Style dropdown. Saved Azure styles on dubbed events are ignored at render time.

In the Virtual Voice Manager, available styles are shown as tags beneath the voice name. Once you have assigned a Microsoft voice to an Audio Description event, the style options for that voice appear as a dropdown menu within the event editor, positioned above the notes field. Selecting a style does not immediately change the stored audio β€” the event must be re-rendered before the style takes effect. The render button flashes green after a style change to indicate that a new render is needed.

Because voice style is stored per-event, different events within the same AD group can use different styles. This can be useful when a project includes both a straightforward narration track and moments where a warmer or more dramatic delivery is more appropriate.