# Animated captions with FFmpeg in Node.js

> Burn word-by-word animated captions into an MP4 from Node.js: no headless browser, no ASS files, the same renderer as the browser preview.

`@captionsjs/server` draws captions.js frames with skia-canvas and pipes them straight
into FFmpeg. You get the kinetic, word-level captions people know from TikTok and Reels,
rendered on your server, with output that matches the browser preview frame for frame.

## Why not drawtext or ASS?

| | captions.js | FFmpeg `drawtext` | ASS / libass |
| --- | --- | --- | --- |
| Word-by-word highlight | yes | build it from many filters | karaoke tags only |
| Per-word scale, bounce, box animations | yes | no | no |
| Same look as a live browser preview | yes, same `renderFrame` | no | depends on the player |
| Headless Chrome needed | no | no | no |

## Install

```bash
npm install @captionsjs/server
```

You need `ffmpeg` and `ffprobe` on `PATH` (or point `FFMPEG_PATH` / `FFPROBE_PATH` at them).
On macOS: `brew install ffmpeg`. On Debian/Ubuntu: `apt-get install ffmpeg`.

## Burn captions

```ts
import { burnCaptions } from "@captionsjs/server";

const result = await burnCaptions({
  video: "talk.mp4",            // path or http(s) URL
  captions: "words.json",       // path, URL, JSON string or parsed array
  preset: "Focus Box",          // any preset name, case-insensitive
  output: "talk.captions.mp4",  // default: <input>.captions.mp4
  onProgress: ({ frame, totalFrames }) =>
    process.stdout.write(`\r${Math.round((frame / totalFrames) * 100)}%`),
});

console.log(result); // { output, frames, width, height, fps, durationSec }
```

`captions` accepts plain `[{ word, start, end }]` (or `startTime` / `endTime`), OpenAI
Whisper `verbose_json` and Deepgram responses. See
[Animated captions from Whisper](/docs/guides/whisper-animated-captions) for getting word timings.

## Options

| Option | Default | What it does |
| --- | --- | --- |
| `preset` | `"Karaoke"` | Preset name or a full preset object (clone one with `getPreset()` and edit it) |
| `fps` | source fps | Frame rate of the caption overlay |
| `crf` | `20` | x264 quality, lower is better and bigger |
| `x264Preset` | `"veryfast"` | x264 speed/size trade-off |
| `scale` | `height / 480` | Font-size multiplier; the browser overlay uses the same rule |
| `fonts.fontsDir` | none | Folder with TTFs in Google Fonts layout, searched first |
| `fonts.download` | `true` | Set `false` to fail instead of downloading missing fonts |

Preset fonts are downloaded from Google Fonts on first use and cached in
`~/.cache/captionsjs/fonts` (set `CAPTIONSJS_CACHE_DIR` to move it).

## Custom styles

A preset is a plain object, so brand it in code and pass it in:

```ts
import { getPreset } from "captions.js";
import { burnCaptions } from "@captionsjs/server";

const base = getPreset("Karaoke");
const branded = {
  ...base,
  captionsSettings: {
    ...base.captionsSettings,
    style: { ...base.captionsSettings.style, highlightColor: "#ff2e88" },
  },
};

await burnCaptions({ video: "talk.mp4", captions: "words.json", preset: branded });
```

## Without writing code

```bash
npx captions.js burn talk.mp4 words.json --preset Karaoke -o out.mp4
npx captions.js presets   # list every preset name
```

Or run the Docker image, which has FFmpeg and every preset font baked in and works offline:

```bash
docker run --rm -v "$PWD:/data" maskin25/captions.js-render \
  burn /data/talk.mp4 /data/words.json --preset Karaoke -o /data/out.mp4
```

More on the image in [Docker Render Service](/docs/server).

## Rather not run it yourself?

[captions.js Cloud](/cloud) is the same renderer as a hosted API: send a video URL, get the
captioned MP4 back by webhook. It's in early access; founding members get free render minutes.

Source: https://captionsjs.dev/docs/guides/ffmpeg-animated-captions-nodejs
