The pipeline

Stage 04 · ElevenLabs-class TTS

A narration booth, driven by the timecode.

Narration is performed against the blueprint's beat map, then delivered with the word-level timing the edit needs for captions, cut points and music ducking.

Delivery

48 kHz, broadcast loudness

Timing data

Word-level timestamps

Pace target

Beat-matched, 140–165 wpm

Loudness

−14 LUFS integrated

01

Performance direction, not just text-to-speech

Each beat carries a direction: pace, emphasis, pause length before a reveal, and the drop in energy that makes a reversal land. Those directions are compiled into the synthesis request, so the read has structure rather than uniform pleasantness.

  • Held pauses before evidence reveals
  • Emphasis marks on the beat's operative noun
  • Breath and room tone preserved between paragraphs
02

Timing exported for the edit

Word-level timestamps come back with the audio. Cut points snap to phrase boundaries, captions are frame-accurate without transcription, and the motion stage knows exactly how long each shot must hold on screen.

03

Mix delivered to spec

The stem arrives loudness-normalised to platform target with de-essing and light compression applied, so the master needs no corrective pass before upload.