← Back to lab

I Replaced a 702-Line ffmpeg Prompt With Remotion for AI-Generated Shorts. Here's the Diff.

My agents built shorts with ffmpeg drawtext at 576p and 134 kbps. Remotion plus one spec.json now renders 1080p reels and 16:9 videos. Numbers, spec, costs.

I Replaced a 702-Line ffmpeg Prompt With Remotion for AI-Generated Shorts. Here's the Diff.

576p, DejaVu Sans, 134 kbps

The old short-video skill was a 702-line prompt teaching an agent to hand-write ffmpeg filter chains. It rendered 576×1024 at 24 fps, burned captions with drawtext in DejaVu Sans Bold, and did "motion" with zoompan. The 47 scene clips still sitting in the temp directory are 1080×1920 at a median bitrate of 134 kbps. They look exactly as cheap as that number suggests.

Model: Claude (agent) + Remotion 4.0.434  |  TTS cost: $0.015 per reel + long video  |  Date: 2026-09  |  Status: in production

The replacement is a Remotion project: React components, one spec.json per video, and a CLI that does TTS, checks, stills, render and QA. The old skill was archived on 2026-09-27 with a one-line reason: "ffmpeg drawtext + zoompan at 576×1024, DejaVu font". The routing skill now opens with a hard rule: never hand-roll PIL frames or ffmpeg drawtext slideshows.

What the agent was actually writing

The core of the old pipeline, condensed from the archived skill:

# One scene. FRAMES = DURATION * 24. Multiply by 5-8 scenes, then concat, then mix audio.
ffmpeg -y -loop 1 -i scene_N.png -t {DURATION} \
  -vf "scale=640:1138,zoompan=z='min(zoom+0.001,1.08)':x='iw/2-(iw/zoom/2)':y='ih/2-(ih/zoom/2)':d={FRAMES}:s=576x1024:fps=24,\
drawtext=textfile=text_Na.txt:fontsize=38:fontcolor=yellow:borderw=3:bordercolor=black:x=(w-text_w)/2:y=h*0.08:fontfile=$FONT" \
  -c:v libx264 -preset fast -crf 23 -an scene_N_final.mp4

Every line of that is a place for an agent to fail:

  • Escaping. The skill carried a bolded warning: use textfile= because inline text= breaks on special characters. Czech has plenty of them. Colons, quotes and percent signs in a headline break the filter graph too.
  • Timing. The agent computed {DURATION} and {FRAMES} per scene. Captions had no notion of time inside a scene: one static overlay per clip.
  • Layout. x=(w-text_w)/2:y=h*0.08 is the whole layout engine. No wrapping, no safe zones, no measuring. Long lines ran off screen.
  • One format. 9:16 only. A 16:9 version meant a second, different set of filter chains.

The agent never looked at what it made. Nothing in the skill told it to.

The new shape: data in, video out

A tutorial video is now addressed as <site>/<slug> and lives in one folder:

public-tutorial/<site>/<slug>/
  spec.json            scenes for both formats
  shots/*.png          real, cropped screenshots
  vo/*.mp3             one file per scene
  vo/durations.json    measured VO lengths, drives the timeline
out/tutorial/<site>/<slug>/
  <out>.mp4, stills/, *-contact.png

The spec is the only thing the agent edits. Sanitized, trimmed excerpt of a real one (the first scene of a 7-scene reel):

{
  "brand": "demo",
  "title": ["Kartičky", "v ChatGPT"],
  "site": "example.com",
  "tts": "openrouter",
  "voices": {
    "K": { "voice": "coral", "role": "host" },
    "T": { "voice": "cedar", "role": "co-host" }
  },
  "tempo": { "short": 1.05, "long": 1.0 },
  "out": { "short": "demo-flashcards", "long": "demo-flashcards-youtube" },
  "short": [
    {
      "id": "s01",
      "kind": "split",
      "headline": ["Stejný prompt, +1 slovo", "obrázek, nebo tohle?"],
      "text": "Stačí jedno slovo v promptu… a místo obrázku máš tohle! Počkej, fakt?",
      "shot":  { "file": "obrazek.png", "w": 940, "h": 660 },
      "shot2": { "file": "zadni.png",   "w": 960, "h": 620 },
      "stamp": { "t": "✗ JEN OBRÁZEK", "at": 0.15 },
      "lines": [
        { "who": "K", "emotion": "napínavě, pak nadšeně odhalíš",
          "say": "Stačí jedno slovo v promptu… a místo obrázku máš tohle!" },
        { "who": "T", "emotion": "překvapeně, nevěřícně", "say": "Počkej, fakt?" }
      ]
    }
  ],
  "long": []
}

Diacritics are just JSON strings rendered by Chromium with a real font. No escaping layer exists to get wrong.

Timing is not something the agent chooses. The TTS step writes durations.json, and the timeline is pure math over it:

// Scene length = measured VO + fixed padding. Never hardcoded.
export const LEAD = 6;  // frames before each VO line
export const TAIL = 8;  // frames after it

const vo = Math.ceil(d * FPS);
const len = LEAD + vo + TAIL;

Everything positional inside a scene (rings, badges, stamp, shot2 slide-in, zoom) takes at as a fraction of that scene's VO. Rewrite a line, re-run TTS, and every highlight moves with the voice. The agent never touches a frame number.

One spec, two formats

The same spec holds a short array (1080×1920) and a long array (1920×1080). Two compositions, TutorialShort and TutorialLong, read it. Measured on the three tutorial videos in production:

VideoShort scenesLong scenesShortLongTTS
Charts tutorial61238.3 s, 5.3 MB2:19, 14.9 MBedge-tts (free)
Subscription explainer71038.3 s, 6.4 MB1:53, 12.5 MBtwo-voice, $0.015 total
Flashcards remake7032.7 s, 5.8 MBnonetwo-voice
All renders: H.264, CRF 18, 30 fps. The shorts land at 1.1–1.4 Mbps, roughly 8–10x the bitrate of the old ffmpeg clips at the same frame size.
## Captions: less magic than advertised
The karaoke captions are not aligned to TTS word timestamps. I assumed they were. They aren't, and that's deliberate.
The daily news short pipeline splits the narration's measured duration evenly across the script's words:
```python
# Uses script text directly, never whisper transcription.
# Czech and English loanwords get mangled by whisper.
ms_per_word = duration_ms / len(all_words)
`
A Remotion component then pages through that SRT four words at a time and highlights the active one. The tutorial captions are cruder still: ~22-character chunks timed by length. Both drift slightly on long words and pauses. Both are still better than whisper's timestamps paired with whisper's spelling of Czech tech jargon. Correct text beats precise timing.
Whisper does have a job, just not a rendering one. It runs in QA, after the render.
## QA the agent can actually do
```bash
python3 scripts/tutorial.py tts $V # only changed lines regenerate
python3 scripts/tutorial.py check $V # files, sizes, rings in bounds
python3 scripts/tutorial.py render $V both --still # stills at 75% of each scene + contact sheets
python3 scripts/tutorial.py render $V both # MP4s, CRF 18
python3 scripts/tutorial.py qa $V short --whisper
`
qa fails on an empty first frame or a length far off target, and warns on pauses over 0.35 s inside a VO line or a whisper-vs-script match below 0.75. The contact sheet is one PNG with a still from every scene. That's the step that changed the output the most: the agent opens one image and sees overlaps, a ring off its target, or a missing diacritic before spending a full render.
On top of that sits a separate read-only video-critic subagent. It gets the whisper transcript with timestamps, the contact sheet, and stats (length, silences, loudness). It scores hook, first frame, pace, readability, voice, CTA and audio 1–5 with evidence, then maps its top 5 fixes to spec changes (kind: split, headline, reorder, cut). The flashcards remake came out of that loop: the split hook, per-scene headlines and the two-voice dialogue are all critic-driven spec edits.
## The diff
ffmpeg / PILRemotion + spec.json
---------
Resolution576×1024, 24 fps1080×1920 and 1920×1080, 30 fps
Bitrate (9:16)median 134 kbps (47 clips)1.1–1.4 Mbps
Captionsone static drawtext per clippaged, per-word highlight, timed from real VO
Timingagent computes framesderived from durations.json
Fonts / diacriticsDejaVu, textfile= workaroundbrand fonts, plain JSON strings
Formats9:16 onlyboth from one spec
Agent editsfilter-graph stringstyped JSON fields
Self-reviewnonestills, contact sheet, qa, critic agent
API cost~$0.04 per short (screenshot mode)$0 edge-tts, ~$0.015 two-voice dialogue
Render costseconds of CPU~1x realtime at 1080p (project note, not re-measured)

What Remotion costs you

Render time. The project notes put it at about 1x realtime at 1080p on the VPS (a 45 s video in ~46 s). I didn't re-measure it for this post. ffmpeg on a 576p slideshow is effectively instant by comparison.

Chromium. Remotion renders through a headless Chrome shell: 207 MB inside node_modules/.remotion, in an 866 MB node_modules.

Bundles. Every render writes a webpack bundle to $TMPDIR, and the bundle copies the public dir. The docs estimate ~70 MB. With a brand-assets folder of 2.6 GB, each bundle is 2.6 GB. Three were left over from renders one minute apart today: 7.8 GB of temp. The tutorial CLI deletes them. Ad-hoc renders don't, and an ENOSPC mid-render is the symptom.

React footguns. Hooks after an early return crash the render. <img> instead of Remotion's <Img> renders frames before the image loads. Agents hit both.

A chatty TTS model. The two-voice dialogue uses an audio chat model. Without a wrapper it answers the line instead of reading it ("Hi!" gets "Hi! How can I help?"). The TTS step retries up to 5x until the model's own transcript matches the script at ≥0.92, caches every line, and exits non-zero on a mismatch.

What I'd change

Real word-level alignment, done by forced alignment of the known script against the audio rather than free transcription, would fix caption drift without letting whisper's spelling back in. I'd also move the bundle's public dir to a per-video subset. Copying 2.6 GB of unrelated brand assets to render a 38-second tutorial is the one thing in this pipeline that's dumber than the ffmpeg version it replaced.