# Animated captions from Whisper word timestamps

> Turn OpenAI Whisper, faster-whisper or the Whisper API output into word-by-word animated captions, in the browser and burned into an MP4.

Whisper gives you **when each word was spoken**. captions.js turns those timings into
animated, word-by-word captions: live over a `<video>` in the browser, or burned into
an MP4 with FFmpeg. Both paths use the same renderer, so the preview and the export match.

## 1. Get word-level timestamps

Captions need a start and end time **per word**, not per segment. Turn word timestamps on:

**OpenAI API** (`whisper-1`): ask for `verbose_json` with word granularity. The words
come back in a top-level `words` array.

```ts
import OpenAI from "openai";
import fs from "node:fs";

const openai = new OpenAI();
const transcript = await openai.audio.transcriptions.create({
  file: fs.createReadStream("talk.mp4"),
  model: "whisper-1",
  response_format: "verbose_json",
  timestamp_granularities: ["word"],
});

fs.writeFileSync("words.json", JSON.stringify(transcript));
// transcript.words → [{ word: "Hello", start: 0.12, end: 0.48 }, …]
```

**openai-whisper** (Python): pass `word_timestamps=True`. The words sit inside
`segments[].words`.

```python
import json, whisper

model = whisper.load_model("small")
result = model.transcribe("talk.mp4", word_timestamps=True)
json.dump(result, open("words.json", "w"))
```

**faster-whisper**: also `word_timestamps=True`; collect the words yourself:

```python
from faster_whisper import WhisperModel
import json

segments, _ = WhisperModel("small").transcribe("talk.mp4", word_timestamps=True)
words = [{"word": w.word.strip(), "start": w.start, "end": w.end}
         for s in segments for w in s.words]
json.dump(words, open("words.json", "w"))
```

## 2. Burn the captions into the video

The server package reads all three shapes above as they are: a top-level `words`
array, `segments[].words`, or a plain `[{ word, start, end }]` list.

```bash
npx captions.js burn talk.mp4 words.json --preset Karaoke
# → talk.captions.mp4
```

Or from Node:

```ts
import { burnCaptions } from "@captionsjs/server";

await burnCaptions({
  video: "talk.mp4",
  captions: "words.json",
  preset: "Karaoke",
  output: "talk.captions.mp4",
});
```

Needs `ffmpeg` and `ffprobe` on `PATH`. See [Animated captions with FFmpeg in Node.js](/docs/guides/ffmpeg-animated-captions-nodejs) for every option.

## 3. Show them live in the browser

In the browser, `toCaptions()` reads a Deepgram response or an array of
`{ word, startTime, endTime }`. Map the Whisper words first:

```ts
import captionsjs, { getPreset } from "captions.js";

type WhisperWord = { word: string; start: number; end: number };

// OpenAI API: transcript.words; openai-whisper: segments.flatMap(s => s.words)
function fromWhisper(words: WhisperWord[]) {
  return words.map((w) => ({ word: w.word.trim(), startTime: w.start, endTime: w.end }));
}

const instance = captionsjs({
  video: document.querySelector("video")!,
  preset: getPreset("Karaoke"),
  captions: fromWhisper(transcript.words),
});
```

The overlay follows `video.currentTime`, so seeking, pausing and playback speed just work.
Switch the look without remounting: `instance.preset(getPreset("Focus Box"))`.

## Tips

- **Pick a style** from the [caption styles gallery](/styles): every preset name works in
  `getPreset()`, `--preset` and `burnCaptions({ preset })`.
- **Clean words first.** Whisper sometimes returns leading spaces or punctuation-only
  tokens; trim them and drop empty words before rendering.
- **Long videos:** transcribe once, save `words.json`, and reuse it for every preview and export.

Source: https://captionsjs.dev/docs/guides/whisper-animated-captions
