captions.js

Animated captions from Whisper word timestamps

Turn OpenAI Whisper, faster-whisper or the Whisper API output into word-by-word animated captions, in the browser and burned into an MP4.

Whisper gives you when each word was spoken. captions.js turns those timings into animated, word-by-word captions: live over a <video> in the browser, or burned into an MP4 with FFmpeg. Both paths use the same renderer, so the preview and the export match.

1. Get word-level timestamps

Captions need a start and end time per word, not per segment. Turn word timestamps on:

OpenAI API (whisper-1): ask for verbose_json with word granularity. The words come back in a top-level words array.

import OpenAI from "openai";
import fs from "node:fs";

const openai = new OpenAI();
const transcript = await openai.audio.transcriptions.create({
  file: fs.createReadStream("talk.mp4"),
  model: "whisper-1",
  response_format: "verbose_json",
  timestamp_granularities: ["word"],
});

fs.writeFileSync("words.json", JSON.stringify(transcript));
// transcript.words → [{ word: "Hello", start: 0.12, end: 0.48 }, …]

openai-whisper (Python): pass word_timestamps=True. The words sit inside segments[].words.

import json, whisper

model = whisper.load_model("small")
result = model.transcribe("talk.mp4", word_timestamps=True)
json.dump(result, open("words.json", "w"))

faster-whisper: also word_timestamps=True; collect the words yourself:

from faster_whisper import WhisperModel
import json

segments, _ = WhisperModel("small").transcribe("talk.mp4", word_timestamps=True)
words = [{"word": w.word.strip(), "start": w.start, "end": w.end}
         for s in segments for w in s.words]
json.dump(words, open("words.json", "w"))

2. Burn the captions into the video

The server package reads all three shapes above as they are: a top-level words array, segments[].words, or a plain [{ word, start, end }] list.

npx captions.js burn talk.mp4 words.json --preset Karaoke
# → talk.captions.mp4

Or from Node:

import { burnCaptions } from "@captionsjs/server";

await burnCaptions({
  video: "talk.mp4",
  captions: "words.json",
  preset: "Karaoke",
  output: "talk.captions.mp4",
});

Needs ffmpeg and ffprobe on PATH. See Animated captions with FFmpeg in Node.js for every option.

3. Show them live in the browser

In the browser, toCaptions() reads a Deepgram response or an array of { word, startTime, endTime }. Map the Whisper words first:

import captionsjs, { getPreset } from "captions.js";

type WhisperWord = { word: string; start: number; end: number };

// OpenAI API: transcript.words; openai-whisper: segments.flatMap(s => s.words)
function fromWhisper(words: WhisperWord[]) {
  return words.map((w) => ({ word: w.word.trim(), startTime: w.start, endTime: w.end }));
}

const instance = captionsjs({
  video: document.querySelector("video")!,
  preset: getPreset("Karaoke"),
  captions: fromWhisper(transcript.words),
});

The overlay follows video.currentTime, so seeking, pausing and playback speed just work. Switch the look without remounting: instance.preset(getPreset("Focus Box")).

Tips

  • Pick a style from the caption styles gallery: every preset name works in getPreset(), --preset and burnCaptions({ preset }).
  • Clean words first. Whisper sometimes returns leading spaces or punctuation-only tokens; trim them and drop empty words before rendering.
  • Long videos: transcribe once, save words.json, and reuse it for every preview and export.

For AI agents and LLMs: this page as Markdown · llms.txt · all docs in one file