576p, DejaVu Sans, 134 kbps
The old short-video skill was a 702-line prompt teaching an agent to hand-write ffmpeg filter chains. It rendered 576×1024 at 24 fps, burned captions with drawtext in DejaVu Sans Bold, and did "motion" with zoompan. The 47 scene clips still sitting in the temp directory are 1080×1920 at a median bitrate of 134 kbps. They look exactly as cheap as that number suggests.
Model: Claude (agent) + Remotion 4.0.434 | TTS cost: $0.015 per reel + long video | Date: 2026-09 | Status: in production The replacement is a Remotion project: React components, one spec.json per video, and a CLI that does TTS, checks, stills, render and QA. The old skill was archived on 2026-09-27 with a one-line reason: "ffmpeg drawtext + zoompan at 576×1024, DejaVu font". The routing skill now opens with a hard rule: never hand-roll PIL frames or ffmpeg drawtext slideshows.
What the agent was actually writing
The core of the old pipeline, condensed from the archived skill:
# One scene. FRAMES = DURATION * 24. Multiply by 5-8 scenes, then concat, then mix audio.
ffmpeg -y -loop 1 -i scene_N.png -t {DURATION} \
-vf "scale=640:1138,zoompan=z='min(zoom+0.001,1.08)':x='iw/2-(iw/zoom/2)':y='ih/2-(ih/zoom/2)':d={FRAMES}:s=576x1024:fps=24,\
drawtext=textfile=text_Na.txt:fontsize=38:fontcolor=yellow:borderw=3:bordercolor=black:x=(w-text_w)/2:y=h*0.08:fontfile=$FONT" \
-c:v libx264 -preset fast -crf 23 -an scene_N_final.mp4 Every line of that is a place for an agent to fail:
- Escaping. The skill carried a bolded warning: use
textfile=because inlinetext=breaks on special characters. Czech has plenty of them. Colons, quotes and percent signs in a headline break the filter graph too. - Timing. The agent computed
{DURATION}and{FRAMES}per scene. Captions had no notion of time inside a scene: one static overlay per clip. - Layout.
x=(w-text_w)/2:y=h*0.08is the whole layout engine. No wrapping, no safe zones, no measuring. Long lines ran off screen. - One format. 9:16 only. A 16:9 version meant a second, different set of filter chains.
The agent never looked at what it made. Nothing in the skill told it to.
The new shape: data in, video out
A tutorial video is now addressed as <site>/<slug> and lives in one folder:
public-tutorial/<site>/<slug>/
spec.json scenes for both formats
shots/*.png real, cropped screenshots
vo/*.mp3 one file per scene
vo/durations.json measured VO lengths, drives the timeline
out/tutorial/<site>/<slug>/
<out>.mp4, stills/, *-contact.png The spec is the only thing the agent edits. Sanitized, trimmed excerpt of a real one (the first scene of a 7-scene reel):
{
"brand": "demo",
"title": ["Kartičky", "v ChatGPT"],
"site": "example.com",
"tts": "openrouter",
"voices": {
"K": { "voice": "coral", "role": "host" },
"T": { "voice": "cedar", "role": "co-host" }
},
"tempo": { "short": 1.05, "long": 1.0 },
"out": { "short": "demo-flashcards", "long": "demo-flashcards-youtube" },
"short": [
{
"id": "s01",
"kind": "split",
"headline": ["Stejný prompt, +1 slovo", "obrázek, nebo tohle?"],
"text": "Stačí jedno slovo v promptu… a místo obrázku máš tohle! Počkej, fakt?",
"shot": { "file": "obrazek.png", "w": 940, "h": 660 },
"shot2": { "file": "zadni.png", "w": 960, "h": 620 },
"stamp": { "t": "✗ JEN OBRÁZEK", "at": 0.15 },
"lines": [
{ "who": "K", "emotion": "napínavě, pak nadšeně odhalíš",
"say": "Stačí jedno slovo v promptu… a místo obrázku máš tohle!" },
{ "who": "T", "emotion": "překvapeně, nevěřícně", "say": "Počkej, fakt?" }
]
}
],
"long": []
} Diacritics are just JSON strings rendered by Chromium with a real font. No escaping layer exists to get wrong.
Timing is not something the agent chooses. The TTS step writes durations.json, and the timeline is pure math over it:
// Scene length = measured VO + fixed padding. Never hardcoded.
export const LEAD = 6; // frames before each VO line
export const TAIL = 8; // frames after it
const vo = Math.ceil(d * FPS);
const len = LEAD + vo + TAIL; Everything positional inside a scene (rings, badges, stamp, shot2 slide-in, zoom) takes at as a fraction of that scene's VO. Rewrite a line, re-run TTS, and every highlight moves with the voice. The agent never touches a frame number.
One spec, two formats
The same spec holds a short array (1080×1920) and a long array (1920×1080). Two compositions, TutorialShort and TutorialLong, read it. Measured on the three tutorial videos in production:
| Video | Short scenes | Long scenes | Short | Long | TTS |
| Charts tutorial | 6 | 12 | 38.3 s, 5.3 MB | 2:19, 14.9 MB | edge-tts (free) |
| Subscription explainer | 7 | 10 | 38.3 s, 6.4 MB | 1:53, 12.5 MB | two-voice, $0.015 total |
| Flashcards remake | 7 | 0 | 32.7 s, 5.8 MB | none | two-voice |
| All renders: H.264, CRF 18, 30 fps. The shorts land at 1.1–1.4 Mbps, roughly 8–10x the bitrate of the old ffmpeg clips at the same frame size. | |||||
| ## Captions: less magic than advertised | |||||
| The karaoke captions are not aligned to TTS word timestamps. I assumed they were. They aren't, and that's deliberate. | |||||
| The daily news short pipeline splits the narration's measured duration evenly across the script's words: | |||||
| ```python | |||||
| # Uses script text directly, never whisper transcription. | |||||
| # Czech and English loanwords get mangled by whisper. | |||||
| ms_per_word = duration_ms / len(all_words) | |||||
| ` | |||||
| A Remotion component then pages through that SRT four words at a time and highlights the active one. The tutorial captions are cruder still: ~22-character chunks timed by length. Both drift slightly on long words and pauses. Both are still better than whisper's timestamps paired with whisper's spelling of Czech tech jargon. Correct text beats precise timing. | |||||
| Whisper does have a job, just not a rendering one. It runs in QA, after the render. | |||||
| ## QA the agent can actually do | |||||
| ```bash | |||||
| python3 scripts/tutorial.py tts $V # only changed lines regenerate | |||||
| python3 scripts/tutorial.py check $V # files, sizes, rings in bounds | |||||
| python3 scripts/tutorial.py render $V both --still # stills at 75% of each scene + contact sheets | |||||
| python3 scripts/tutorial.py render $V both # MP4s, CRF 18 | |||||
| python3 scripts/tutorial.py qa $V short --whisper | |||||
| ` | |||||
| qa fails on an empty first frame or a length far off target, and warns on pauses over 0.35 s inside a VO line or a whisper-vs-script match below 0.75. The contact sheet is one PNG with a still from every scene. That's the step that changed the output the most: the agent opens one image and sees overlaps, a ring off its target, or a missing diacritic before spending a full render. | |||||
| On top of that sits a separate read-only video-critic subagent. It gets the whisper transcript with timestamps, the contact sheet, and stats (length, silences, loudness). It scores hook, first frame, pace, readability, voice, CTA and audio 1–5 with evidence, then maps its top 5 fixes to spec changes (kind: split, headline, reorder, cut). The flashcards remake came out of that loop: the split hook, per-scene headlines and the two-voice dialogue are all critic-driven spec edits. | |||||
| ## The diff | |||||
| ffmpeg / PIL | Remotion + spec.json | ||||
| --- | --- | --- | |||
| Resolution | 576×1024, 24 fps | 1080×1920 and 1920×1080, 30 fps | |||
| Bitrate (9:16) | median 134 kbps (47 clips) | 1.1–1.4 Mbps | |||
| Captions | one static drawtext per clip | paged, per-word highlight, timed from real VO | |||
| Timing | agent computes frames | derived from durations.json | |||
| Fonts / diacritics | DejaVu, textfile= workaround | brand fonts, plain JSON strings | |||
| Formats | 9:16 only | both from one spec | |||
| Agent edits | filter-graph strings | typed JSON fields | |||
| Self-review | none | stills, contact sheet, qa, critic agent | |||
| API cost | ~$0.04 per short (screenshot mode) | $0 edge-tts, ~$0.015 two-voice dialogue | |||
| Render cost | seconds of CPU | ~1x realtime at 1080p (project note, not re-measured) |
What Remotion costs you
Render time. The project notes put it at about 1x realtime at 1080p on the VPS (a 45 s video in ~46 s). I didn't re-measure it for this post. ffmpeg on a 576p slideshow is effectively instant by comparison.
Chromium. Remotion renders through a headless Chrome shell: 207 MB inside node_modules/.remotion, in an 866 MB node_modules.
Bundles. Every render writes a webpack bundle to $TMPDIR, and the bundle copies the public dir. The docs estimate ~70 MB. With a brand-assets folder of 2.6 GB, each bundle is 2.6 GB. Three were left over from renders one minute apart today: 7.8 GB of temp. The tutorial CLI deletes them. Ad-hoc renders don't, and an ENOSPC mid-render is the symptom.
React footguns. Hooks after an early return crash the render. <img> instead of Remotion's <Img> renders frames before the image loads. Agents hit both.
A chatty TTS model. The two-voice dialogue uses an audio chat model. Without a wrapper it answers the line instead of reading it ("Hi!" gets "Hi! How can I help?"). The TTS step retries up to 5x until the model's own transcript matches the script at ≥0.92, caches every line, and exits non-zero on a mismatch.
What I'd change
Real word-level alignment, done by forced alignment of the known script against the audio rather than free transcription, would fix caption drift without letting whisper's spelling back in. I'd also move the bundle's public dir to a per-video subset. Copying 2.6 GB of unrelated brand assets to render a 38-second tutorial is the one thing in this pipeline that's dumber than the ffmpeg version it replaced.