← Back to lab

My Disk Cleanup Deleted the Only Browser Build Three Pipelines Needed. Twice.

A 'keep newest per family' cleanup rule deleted the exact Playwright build my Python scrapers pin. Three pipelines broke at 02:45, and nobody noticed for 22 hours.

My Disk Cleanup Deleted the Only Browser Build Three Pipelines Needed. Twice.

Date: 2026-10 | Root cause: heuristic delete | Blast radius: 3 pipelines, 2 recurrences | Detection lag: ~22 hours

At 02:45 on a Saturday, the visual-QA check on one of my content pipelines stopped working. Not failed — stopped. The job ran, logged an error about a missing executable, and the pipeline moved on without its screenshots. I found out roughly 22 hours later, from a nightly ops summary, not from the pipeline itself.

The executable was a Playwright browser build. My own disk-cleanup script had deleted it six hours earlier.

The offending code

The server runs a nightly cleanup script. One section prunes the Playwright browser cache, which otherwise grows without bound:

# Playwright browsers — keep only newest version of each browser family
PW=~/.cache/ms-playwright
for family in chromium chromium_headless_shell firefox webkit ffmpeg; do
    newest=$(ls -dt "$PW/$family"-*/ 2>/dev/null | head -1 || true)
    [ -n "$newest" ] || continue
    for v in "$PW/$family"-*/; do
        [ "$v" = "$newest" ] && continue
        log "Removing old playwright $family $(basename "$v")"
        rm -rf "$v"
    done
done

The logic reads clean. Old browser builds are dead weight — each chromium revision is 393 MB unpacked (261 MB more for the headless shell), so a cache holding three revisions burns close to 2 GB. Keep the newest, delete the rest, done.

Except "old" and "unused" are not the same thing.

Why newest-per-family is wrong

Playwright doesn't use "a chromium." Each Playwright package version pins an exact browser revision. Python Playwright 1.58 wants chromium 1208. The Node apps on the same box pull 1228. Both live side by side in ~/.cache/ms-playwright/ and both are in active use:

  • A Python-based scraper fallback and a screenshot-research step need 1208.
  • Two Node services need 1228.

So the cache holds two revisions on purpose. The cleanup script looks at that directory, decides 1208 is stale because 1228 is newer, and deletes it. Every Python-driven browser job then dies with Playwright's least helpful error:

browserType.launch: Executable doesn't exist at
~/.cache/ms-playwright/chromium-1208/chrome-linux/chrome
╔══════════════════════════════════════════════╗
║ Looks like Playwright was just updated.      ║
║ Please run: playwright install               ║
╚══════════════════════════════════════════════╝

Playwright was not just updated. The binary was there at midnight and gone by 02:00.

The blast radius

One rm -rf took out three separate things:

  1. The visual-QA check — a daily screenshot diff that watches rendered pages for layout breakage. Dead from 02:45 onward.
  2. The scraper fallback — when the plain HTTP fetcher gets blocked, a headless-browser fallback renders the page. Gone.
  3. Article research screenshots — the hands-on test step that captures real UI for tutorials. Gone.

None of these jobs own the browser cache. None of them log "someone deleted my dependency." They just fail, one by one, over the following day, each looking like its own isolated bug. The only reason the root cause surfaced at all is that a separate nightly report happened to diff the cache directory and noticed a revision had vanished.

And it was the second time. The same deletion broke the same pipeline in July, got "fixed" by reinstalling the browser, and recurred the next time disk pressure triggered the cleanup. Reinstalling the dependency without fixing the deleter is not a fix — it's a scheduled recurrence.

The fix: delete by reference, not by heuristic

Every installed playwright-core ships a browsers.json manifest that says exactly which revisions it needs. The correct cleanup keeps the union of every manifest on the box, plus the newest build per family as slack:

import json, os, shutil
from pathlib import Path

def referenced_revisions(roots):
    """Revisions pinned by any installed playwright-core on the box."""
    revs = set()
    for root in roots:
        for f in Path(root).rglob("browsers.json"):
            if "playwright-core" not in str(f):
                continue
            for b in json.load(open(f))["browsers"]:
                revs.add(f'{b["name"]}-{b["revision"]}')
    return revs

def prune(cache, keep):
    for entry in os.listdir(cache):
        if entry not in keep and not is_newest_of_family(cache, entry):
            shutil.rmtree(os.path.join(cache, entry))

Thirty lines. The script now asks "who references this?" before deleting, instead of "does this look old?"

The alternative fix — force every consumer onto one Playwright version — sounds cleaner and fails in practice. Python and Node release on different cadences, and each runtime resolves its own pinned revision at install time. A shared PLAYWRIGHT_BROWSERS_PATH doesn't merge version requirements; it just moves the collision.

What I found

  • The deletion ran at roughly 01:30. First downstream failure at 02:45. Root cause identified around 22 hours later, by accident, from an unrelated report.
  • The "keep newest" rule had been correct for months — right up until the box gained a second Playwright runtime with a different pinned revision. Heuristics that are safe with one consumer become destructive with two.
  • The recurring nature was invisible in per-pipeline logs. Each failure looked local. Only a cross-cutting view (the nightly summary comparing cache contents day over day) connected 01:30 to 02:45.
Failure modePer-pipeline logCross-cutting view
Cleanup deleted a dependency"Executable doesn't exist"Revision vanished between runs
Pipeline keeps running degradedExit code 0, error in stderrMissing output artifacts
RecurrenceLooks like a new bug each timeSame revision deleted twice

What I'd change

Two things, in order. First, any cleanup that deletes by pattern instead of by reference count is a time bomb — the blast radius always lands on a process that doesn't own the directory and can't see the deletion coming. Second, detection lag is the real cost. 22 hours of degraded pipelines cost more than 2 GB of disk ever would. A one-line check — "does every revision in every installed browsers.json still exist?" — running after the cleanup would have turned a silent day-long outage into an immediate alert. Disk is cheap. Undetected degradation isn't.