Date: 2026-10 | Root cause: heuristic delete | Blast radius: 3 pipelines, 2 recurrences | Detection lag: ~22 hours
At 02:45 on a Saturday, the visual-QA check on one of my content pipelines stopped working. Not failed — stopped. The job ran, logged an error about a missing executable, and the pipeline moved on without its screenshots. I found out roughly 22 hours later, from a nightly ops summary, not from the pipeline itself.
The executable was a Playwright browser build. My own disk-cleanup script had deleted it six hours earlier.
The offending code
The server runs a nightly cleanup script. One section prunes the Playwright browser cache, which otherwise grows without bound:
# Playwright browsers — keep only newest version of each browser family
PW=~/.cache/ms-playwright
for family in chromium chromium_headless_shell firefox webkit ffmpeg; do
newest=$(ls -dt "$PW/$family"-*/ 2>/dev/null | head -1 || true)
[ -n "$newest" ] || continue
for v in "$PW/$family"-*/; do
[ "$v" = "$newest" ] && continue
log "Removing old playwright $family $(basename "$v")"
rm -rf "$v"
done
done The logic reads clean. Old browser builds are dead weight — each chromium revision is 393 MB unpacked (261 MB more for the headless shell), so a cache holding three revisions burns close to 2 GB. Keep the newest, delete the rest, done.
Except "old" and "unused" are not the same thing.
Why newest-per-family is wrong
Playwright doesn't use "a chromium." Each Playwright package version pins an exact browser revision. Python Playwright 1.58 wants chromium 1208. The Node apps on the same box pull 1228. Both live side by side in ~/.cache/ms-playwright/ and both are in active use:
- A Python-based scraper fallback and a screenshot-research step need 1208.
- Two Node services need 1228.
So the cache holds two revisions on purpose. The cleanup script looks at that directory, decides 1208 is stale because 1228 is newer, and deletes it. Every Python-driven browser job then dies with Playwright's least helpful error:
browserType.launch: Executable doesn't exist at
~/.cache/ms-playwright/chromium-1208/chrome-linux/chrome
╔══════════════════════════════════════════════╗
║ Looks like Playwright was just updated. ║
║ Please run: playwright install ║
╚══════════════════════════════════════════════╝ Playwright was not just updated. The binary was there at midnight and gone by 02:00.
The blast radius
One rm -rf took out three separate things:
- The visual-QA check — a daily screenshot diff that watches rendered pages for layout breakage. Dead from 02:45 onward.
- The scraper fallback — when the plain HTTP fetcher gets blocked, a headless-browser fallback renders the page. Gone.
- Article research screenshots — the hands-on test step that captures real UI for tutorials. Gone.
None of these jobs own the browser cache. None of them log "someone deleted my dependency." They just fail, one by one, over the following day, each looking like its own isolated bug. The only reason the root cause surfaced at all is that a separate nightly report happened to diff the cache directory and noticed a revision had vanished.
And it was the second time. The same deletion broke the same pipeline in July, got "fixed" by reinstalling the browser, and recurred the next time disk pressure triggered the cleanup. Reinstalling the dependency without fixing the deleter is not a fix — it's a scheduled recurrence.
The fix: delete by reference, not by heuristic
Every installed playwright-core ships a browsers.json manifest that says exactly which revisions it needs. The correct cleanup keeps the union of every manifest on the box, plus the newest build per family as slack:
import json, os, shutil
from pathlib import Path
def referenced_revisions(roots):
"""Revisions pinned by any installed playwright-core on the box."""
revs = set()
for root in roots:
for f in Path(root).rglob("browsers.json"):
if "playwright-core" not in str(f):
continue
for b in json.load(open(f))["browsers"]:
revs.add(f'{b["name"]}-{b["revision"]}')
return revs
def prune(cache, keep):
for entry in os.listdir(cache):
if entry not in keep and not is_newest_of_family(cache, entry):
shutil.rmtree(os.path.join(cache, entry)) Thirty lines. The script now asks "who references this?" before deleting, instead of "does this look old?"
The alternative fix — force every consumer onto one Playwright version — sounds cleaner and fails in practice. Python and Node release on different cadences, and each runtime resolves its own pinned revision at install time. A shared PLAYWRIGHT_BROWSERS_PATH doesn't merge version requirements; it just moves the collision.
What I found
- The deletion ran at roughly 01:30. First downstream failure at 02:45. Root cause identified around 22 hours later, by accident, from an unrelated report.
- The "keep newest" rule had been correct for months — right up until the box gained a second Playwright runtime with a different pinned revision. Heuristics that are safe with one consumer become destructive with two.
- The recurring nature was invisible in per-pipeline logs. Each failure looked local. Only a cross-cutting view (the nightly summary comparing cache contents day over day) connected 01:30 to 02:45.
| Failure mode | Per-pipeline log | Cross-cutting view |
| Cleanup deleted a dependency | "Executable doesn't exist" | Revision vanished between runs |
| Pipeline keeps running degraded | Exit code 0, error in stderr | Missing output artifacts |
| Recurrence | Looks like a new bug each time | Same revision deleted twice |
What I'd change
Two things, in order. First, any cleanup that deletes by pattern instead of by reference count is a time bomb — the blast radius always lands on a process that doesn't own the directory and can't see the deletion coming. Second, detection lag is the real cost. 22 hours of degraded pipelines cost more than 2 GB of disk ever would. A one-line check — "does every revision in every installed browsers.json still exist?" — running after the cleanup would have turned a silent day-long outage into an immediate alert. Disk is cheap. Undetected degradation isn't.