← Back to lab

One Sanitizer Default Blanked 17 Articles for 84 Days. Zero Errors Logged.

Symfony HtmlSanitizer caps input at 20,000 chars and returns an empty string past it. 17 of my pages rendered blank for 84 days, no error anywhere. The check that caught it.

One Sanitizer Default Blanked 17 Articles for 84 Days. Zero Errors Logged.

# One Sanitizer Default Blanked 17 Articles for 84 Days. Zero Errors Logged.

Symfony's HtmlSanitizer ships with maxInputLength = 20000. When your input is longer, sanitize() doesn't throw, doesn't truncate, doesn't log. It returns an empty string. On one of my content sites, that single default blanked the bodies of 17 published articles for 84 days.

Where the sanitizer sits

The site is a conventional server-rendered stack: articles live in the database as markdown with some HTML, a Twig extension renders them to HTML, and before output the rendered string passes through Symfony's HtmlSanitizer. That last hop exists for a good reason — the content is part AI-generated and part user-submitted, and sanitizing on render means the DB never has to be trusted.

The publish path verified everything up to the database: API accepted the article, row exists, response 200. What the page looked like after the sanitizer? Nobody checked. There was no check to fail.

Stack: Symfony + Twig  |  Broken: 2026-07-05 → 2026-09-27 (84 days)  |  Pages affected: 28  |  Status: fixed

What 20,000 characters means in practice

The cap counts characters of rendered HTML, not words. Markdown expands on the way out: every paragraph brings tags, code blocks bring spans, quotes bring entities. A long-form piece crosses 20k characters of HTML without anyone thinking of it as a 20k-character document.

That's the nasty part of the distribution. Short articles render fine. The cap lands exactly on the long-form pieces — the ones with the most research and the most SEO value. 17 published pages rendered a complete page shell: title, navigation, metadata, comments section, and an empty body. HTTP 200 every time.

And the failure mode is the quietest one available:

// Symfony's documented behavior past maxInputLength:
sanitize('...21,000 characters...');  // === ''

No exception. No log line. No truncated output that a reader might notice and report. An empty string, flowing into the template as if the article had no body.

The numbers

MetricValue
Articles with blank bodies17
Articles with body replaced by section summaries11
Days live before the fix84 (2026-07-05 → 2026-09-27)
Error log entries from the failure0
URLs re-pinged to the indexing API after the fix27

The 11 in the second row were a different bug the same sweep found, which is its own lesson. The tutorial template had a "render structured sections instead of the body" branch: whenever a structured-content field was present, readers got the section summaries and the full body never printed. Two independent render paths, two independent ways to silently drop everything below the headline.

The fix

One line on the sanitizer config:

$sanitizer = new HtmlSanitizer(
    (new HtmlSanitizerConfig())
        ->withMaxInputLength(-1)  // default is 20000; past it sanitize() returns ''
        ->allowSafeElements()
);

I don't want the sanitizer enforcing article length at all. Input length is bounded by the editor and the API that writes articles — the render path should render what it's given, and a security filter should filter, not reject-by-emptiness. If you do keep a cap as a resource guard, exceeding it has to be loud: throw, log, alert. Returning '' is a fail-closed door dressed as a success.

The template bug got the same treatment: the body always renders now, structured sections render below it as extras. Never sections XOR body.

The guard: compare rendered words to stored words

The check that catches this class of bug is unglamorous and works on any stack — fetch the rendered page from origin and count the words against the source of truth:

# Origin fetch with a Host header skips the CDN; you're testing your own render,
# not the CDN cache or some edge worker's error page.
curl -s -H "Host: $SITE" "http://127.0.0.1/content/$slug" -o /tmp/page.html
import re, sys, html

raw = open("/tmp/page.html").read()
# strip tags, unescape entities, count non-space runs
text = html.unescape(re.sub(r"<[^>]+>", " ", raw))
print(len(re.findall(r"\S+", text)))

Compare that to the word count of the stored article. A healthy render comes in at the same order of magnitude. A body below ~40% of the stored count means the render path ate the article — blank sanitizer output, body-swapping template, gating logic misfiring, whatever the mechanism is this month.

One exclusion: pages behind an email gate legitimately render a short preview. Check the gate flag before comparing, or compare against the preview length instead of the full body.

What I'd change

The publish pipeline now verifies the render, not the write: after every publish it fetches the page from origin, asserts the title, asserts no raw markdown artifacts leaked through, and asserts the body word count. It would have caught this on day one instead of day 84. A dev-environment assertion would also work — input non-empty and sanitize() output empty should never pass silently — but the origin check is the one that covers the template bug, the gating bug, and the bugs I haven't met yet.