Animated captions from Whisper word timestamps
Turn OpenAI Whisper, faster-whisper or the Whisper API output into word-by-word animated captions, in the browser and burned into an MP4.
Whisper gives you when each word was spoken. captions.js turns those timings into
animated, word-by-word captions: live over a <video> in the browser, or burned into
an MP4 with FFmpeg. Both paths use the same renderer, so the preview and the export match.
1. Get word-level timestamps
Captions need a start and end time per word, not per segment. Turn word timestamps on:
OpenAI API (whisper-1): ask for verbose_json with word granularity. The words
come back in a top-level words array.
import OpenAI from "openai";
import fs from "node:fs";
const openai = new OpenAI();
const transcript = await openai.audio.transcriptions.create({
file: fs.createReadStream("talk.mp4"),
model: "whisper-1",
response_format: "verbose_json",
timestamp_granularities: ["word"],
});
fs.writeFileSync("words.json", JSON.stringify(transcript));
// transcript.words → [{ word: "Hello", start: 0.12, end: 0.48 }, …]openai-whisper (Python): pass word_timestamps=True. The words sit inside
segments[].words.
import json, whisper
model = whisper.load_model("small")
result = model.transcribe("talk.mp4", word_timestamps=True)
json.dump(result, open("words.json", "w"))faster-whisper: also word_timestamps=True; collect the words yourself:
from faster_whisper import WhisperModel
import json
segments, _ = WhisperModel("small").transcribe("talk.mp4", word_timestamps=True)
words = [{"word": w.word.strip(), "start": w.start, "end": w.end}
for s in segments for w in s.words]
json.dump(words, open("words.json", "w"))2. Burn the captions into the video
The server package reads all three shapes above as they are: a top-level words
array, segments[].words, or a plain [{ word, start, end }] list.
npx captions.js burn talk.mp4 words.json --preset Karaoke
# → talk.captions.mp4Or from Node:
import { burnCaptions } from "@captionsjs/server";
await burnCaptions({
video: "talk.mp4",
captions: "words.json",
preset: "Karaoke",
output: "talk.captions.mp4",
});Needs ffmpeg and ffprobe on PATH. See Animated captions with FFmpeg in Node.js for every option.
3. Show them live in the browser
In the browser, toCaptions() reads a Deepgram response or an array of
{ word, startTime, endTime }. Map the Whisper words first:
import captionsjs, { getPreset } from "captions.js";
type WhisperWord = { word: string; start: number; end: number };
// OpenAI API: transcript.words; openai-whisper: segments.flatMap(s => s.words)
function fromWhisper(words: WhisperWord[]) {
return words.map((w) => ({ word: w.word.trim(), startTime: w.start, endTime: w.end }));
}
const instance = captionsjs({
video: document.querySelector("video")!,
preset: getPreset("Karaoke"),
captions: fromWhisper(transcript.words),
});The overlay follows video.currentTime, so seeking, pausing and playback speed just work.
Switch the look without remounting: instance.preset(getPreset("Focus Box")).
Tips
- Pick a style from the caption styles gallery: every preset name works in
getPreset(),--presetandburnCaptions({ preset }). - Clean words first. Whisper sometimes returns leading spaces or punctuation-only tokens; trim them and drop empty words before rendering.
- Long videos: transcribe once, save
words.json, and reuse it for every preview and export.
For AI agents and LLMs: this page as Markdown · llms.txt · all docs in one file