← Back to lab

Approve Buttons Are Rubber Stamps: 8 Weeks of an AI Content Autopilot on Quality Gates

Approve buttons on a scheduled AI content pipeline turned into rubber stamps. Eight weeks of self-enforced quality gates instead: what they caught and missed.

Approve Buttons Are Rubber Stamps: 8 Weeks of an AI Content Autopilot on Quality Gates

Approve Buttons Are Rubber Stamps: 8 Weeks of an AI Content Autopilot on Quality Gates

Window:   2026-08-06 → 2026-09-28  |  Refresh runs: 14  |  Content refreshes: 47
Title/meta fixes: 7  |  Merges (301'd duplicates): 21  |  Approval clicks: 0

The daily article job on a Czech AI-business site I run used to end with a Telegram message: "draft ready, approve?" I tapped approve on almost everything. The one time it mattered, I approved an article about saving money on an LLM API whose model was already a generation old. The hook was stale, the claims were unverified, and it went live with my approval.

That was the moment the approval step got deleted. It was replaced with gates the pipeline enforces on itself, plus one rule: if nothing passes the gates, publish nothing.

Why the button failed

An approve button on a daily job has three failure modes, and I hit all of them:

  • Rubber stamp. Tapping "approve" on a phone gives you a title and maybe a first paragraph. Nobody checks prices, model versions or whether the article duplicates one from last month.
  • Bottleneck. A draft waiting for approval is a draft not indexed. On a news-driven site, a day of latency costs more than most errors.
  • Wrong layer. The bad article was bad at topic selection, before a single word was written. Reviewing output can't fix a bad input.

The fix belonged upstream: make the pipeline refuse bad topics and verify its own output, the way CI refuses a red build.

Architecture

An in-process Node scheduler lives inside a small UI server under PM2, checks a JSON task list every 30 seconds, and spawns headless Claude Code sessions running slash-command skills:

pm2
 └─ ui-server (node)
     └─ scheduler loop (every 30s)
         ├─ daily 08:30      → /article      (write + publish)
         ├─ 2x weekly ~04:00 → /refresh      (update stale, merge duplicates)
         ├─ weekly           → /topics       (rank topic backlog)
         └─ on N consecutive failures → escalate to Telegram

Telegram is still there, but as a log, not a gate. Every run sends a summary. None of them wait for a reply.

The gate chain

This is the order the article and refresh jobs apply checks. Each gate can end the run with "success, nothing published".

gates:
  - name: staleness
    # Kills the exact failure that started this
    reject_if:
      - topic references a model 2+ generations behind current
      - release/announcement news with no concrete workflow angle
      - "model A vs model B price" hook with no how-to-use-it-at-work
    on_all_rejected: publish_nothing   # no fallback "pick anything"

  - name: convergence
    # Topic backlog ranking; the same top picks N runs in a row = stop ranking
    if: same_top_k_for_runs >= 3
    then: recommend_write, stop_rerank

  - name: cannibalization
    # Before writing or refreshing: does this intent already have a page?
    detector: title_similarity >= 75 AND shared_gsc_queries >= 3
    manual_triage:
      - group by model/tool family, not by title
      - GSC query-split: same query, 2+ pages with impressions
      - merge only if same intent AND overlapping tool set
    winner: page with more live inbound internal links

  - name: language
    # Czech site: diacritics, typos, formal/informal register
    check: diacritics preserved, register matches existing body

  - name: render
    # HTTP 200 is not "done"
    check:
      - fetch live page, extract H2s, count words
      - compare against DB body length
      - body has real newlines, not literal "\n"

What the refresh jobs actually caught

Fourteen refresh runs, capped at 3–5 refreshes and at most 3 merges per run.

RunRefreshesTitle/metaMergesNotable catch
08-06303three exact-duplicate pages
08-10212merge winner self-linked to loser slug (redirect loop)
08-20413merge winner carried a stale model-pricing table
08-24322body truncated to 181 words, ending mid-table
08-27210industry stat off by 2x (80% cited, source said 40%)
08-31301AI Act dates corrected to real 2026 deadlines
09-03402third "how much does X cost" duplicate on one topic
09-07303article still called a now-GA feature "a hack"
09-10301merge found only via GSC query-split
09-14401dual annual/monthly pricing reconciled
09-17410free tier pitched as alternative had ended in June
09-21401Czech CRM storage price 2x too high
09-24400body stored with literal \n, zero rendered headings
09-28412last live duplicate of an already-merged cluster
Total47721

The factual corrections are the part an approve button would never have caught:

  • Gemini Code Assist's free and individual tiers ended on 2026-06-18. An AI-coding pricing article was still recommending it as the free option.
  • A Czech CRM's pricing page said 500 Kč per 100 GB of storage. The article said 1000.
  • A scheduled-tasks article claimed a 3-day expiry. Official docs said 7.
  • A product that went GA on June 16 was still described as "limited preview".
  • An article said one open-source coding agent had been renamed to another. It was a fork.

None of these look wrong in a Telegram preview. They look wrong only when a job re-fetches the source and compares.

The render gate earns its keep

Two of the worst bugs returned HTTP 200 and looked fine in the API.

# 08-24: body in DB
words: 181
last line: "| Funkce | Agent mód | Background Agents |"
# ranked at position 43.7 with 77 impressions — Google saw the stub

# 09-24: body in DB
"## Heading\n\nParagraph ..."   # shape of it: literal backslash-n, not newlines
# rendered: one paragraph, 0 headings

A third class is invisible even to a word count: some content types render from a structured field and ignore the body entirely. A body refresh on one of those "succeeds", passes the 200 check and changes nothing on the live page. The gate now checks content type before touching the body and skips those pages.

The lesson came from an admin page earlier: 302 to login was reported as "route works". The first real click returned a 500. Now "done" means the page was fetched and its headings were extracted, not that a status code looked right.

The cannibalization detector returned 0 every time

The similarity detector (title similarity ≥75 plus ≥3 shared GSC queries) confirmed 0 pairs in every run. All 21 merges came from manual triage inside the job:

  • Model-family grouping. Four thin articles about the same model release had different framings, so title similarity stayed under 75. Real cannibalization anyway.
  • Query split. Pulling GSC by query × page and flagging queries where 2+ pages get impressions found a pair (position 5.9 vs 11.3 on the same query) the detector missed.
  • Known false-positive families. On 09-17, 21 of 32 candidates were per-tool pricing pages built from one title template. Same template, different keyword intent. The triage rules now whitelist that cluster.

Fresh duplicates are the worst case: under a month old, too few queries to overlap, invisible to the GSC threshold. Only the family grouping catches them.

Convergence: when to stop ranking

The weekly topics job re-ranks the backlog. By 09-23 the #1 topic had been #1 for 9 consecutive runs, #2 for 8, #3 for 7. The backlog survived 4 weeks of daily publishing (14 new articles per week) without a single one taking a top-3 slot.

That is a signal, not a ranking. More ranking runs add zero information. The gate now says: top-k stable for 3+ runs means write it, stop re-ranking.

Where the human gate stayed

Full newsletter broadcasts. A test send to my own inbox goes out automatically; the send to the full subscriber list waits for an explicit yes, every time.

The logic is asymmetric cost. A bad article can be refreshed, merged or 301'd tomorrow, and the refresh job does exactly that twice a week. A bad email to every verified subscriber can't be recalled and dents sender reputation. Reversible actions get gates. Irreversible ones get a human.

What I'd change

The convergence gate fires but doesn't act. It sends "write #1, stop re-ranking" to Telegram, and the daily job keeps pulling from the news feed instead. Nine runs of the same recommendation is proof that a recommendation is just an approve button with extra steps. The next version routes a converged topic straight into the article job's queue, gated by the same staleness and cannibalization checks as everything else.