← Back to lab

AI Tool Articles Stay Wrong for 46 Days Before Anyone Notices

11 real staleness events from a Czech AI content site: median 46 days between a vendor change and my refresh pipeline catching it. Rot types, cadence, fixes.

AI Tool Articles Stay Wrong for 46 Days Before Anyone Notices

46 days of being wrong, on average

Median lag between a vendor shipping a change and my refresh pipeline fixing the article: 46 days. Fastest catch: 19 days. Slowest: 98. Over that window the article ranks, gets clicks, and states something false with total confidence.

Data:    14 refresh runs, 2026-08-06 → 2026-09-28 (54 days)
Site:    a Czech AI content site I run, ~317 published items
Output:  54 refreshes/title fixes, 21 merges (301s), 0 found by the automated detector
Status:  still running, twice a week

The site covers AI models, coding tools, and SaaS pricing. That is the worst possible niche for content shelf life. Every data point below comes from run logs, not from memory.

The staleness log

Only events where both dates are known: when the fact changed, and when a refresh run caught it.

What changedChangedCaughtLag (days)How it was caught
Agent Skills went GA (article still called skills "a hack")2026-08-192026-09-0719Stale sweep
GPT-5.6 Sol/Terra/Luna shipped, 4 GPT-5.4 articles now obsolete2026-07-092026-07-3021Manual family review
Gemini 3.8 Flash launched (intro $0.75/$3.75)2026-09-022026-09-2422Refresh of the 3.6 Flash article
Platform API breaking change (e-commerce integration guide)2026-08-182026-09-2134Stale sweep
Sonnet 5 $2/$10 intro pricing made permanent2026-08-102026-09-2445Stale sweep; article's core thesis flipped
GPT-5.6 Sol became default for ChatGPT Plus/Pro2026-08-062026-09-2146Stale sweep
MCP 2026-07-28 spec: stateless, init handshake removed2026-07-282026-09-1751Stale sweep
Google Sheets =AI() gained Czech support2026-07-072026-09-0762Stale sweep
Copilot Cowork GA (article said "limited preview")2026-06-162026-08-2469Found during a merge
Gemini Code Assist free/individual tiers ended2026-06-182026-09-1791Stale sweep
Nano Banana Pro GA; DALL-E 3 deprecated2026-05-282026-09-0398Stale sweep
Sorted lags: 19, 21, 22, 34, 45, 46, 51, 62, 69, 91, 98. Median 46, mean 51.
The worst one is the Gemini Code Assist row. The article pitched it as the free alternative to paid coding assistants. For 91 days it recommended a product tier that no longer existed.
Write-to-wrong is shorter than detect-to-fix. A four-way model comparison written on 2026-07-03 went wrong on 2026-08-06 when the ChatGPT default model changed: 34 days of shelf life. The GPT line went 5.4 (March) → 5.5 (April) → 5.6 (July 9). Anything naming "the latest GPT" had a 1–3 month lifespan.
## Five rot types
Grouping every event from the logs by what went wrong:
Rot typeExamples from the logsTypical blast radius
---------
Version churnGPT-5.4→5.5→5.6; Opus 4.6/4.7→Sonnet 5/Opus 4.8; Gemini 3.5/3.6/3.7/3.8 Flash in one quarter; FAQ still saying GPT-4Every article mentioning the model, plus comparison tables
PricingCopilot Pro $10→$15, Canva $12.99→$18, InVideo $17→$25, OpenArt $7→$14, Make €9→$12, ElevenLabs Scale listed at $330, actually $299Every pricing table and "cheapest option" recommendation
Free-tier removalGemini Code Assist free tier (June 2026), a video tool dropping its free plan, n8n Cloud €0→€20Recommendations that were built entirely on "it's free"
DeprecationDALL-E 3 API removed May 2026Tutorials with code that now fails
Feature GA / behavior changeCowork preview→GA, Agent Skills GA, 1M context GA on the Claude 5 family, scheduled task expiry 3 days→7 days, n8n 2.0 moving the task runner to a separate imageArticles framed around "coming soon" or around a limitation that no longer exists

Pricing rot has a nasty variant: the fact itself has more dimensions than the article modeled. The tool database said Copilot Pro was $15. The vendor's pricing page showed $10 on annual billing, $15 on monthly. Both were "right". The article was wrong until it stored both.

Another variant: the tools DB stored a startingPrice field and a separate pricing JSON array for tier detail. Patching one left the other stale, so the generated page contradicted itself. The fix was a sync check:

-- every tool where the generated page disagrees with the source column
SELECT slug, starting_price,
       JSON_EXTRACT(page_content, '$.pricing.starting_price') AS page_price
FROM tools
WHERE JSON_EXTRACT(page_content, '$.pricing.starting_price') <> starting_price;

After a full update run: 55 tools checked, 16 patched, 54 in sync.

New versions create duplicate articles

Model launches do more than make old articles wrong. They make old articles compete with new ones.

When GPT-5.4 shipped, the site produced four thin articles (740–1096 words each): computer use, context window, office tasks, mini. Different headlines, same query intent. The cannibalization detector scores title similarity with a ≥75 threshold plus ≥3 shared Search Console queries. The four scored below 75 and flagged 0 pairs. They were consolidated by hand into one 1640-word article covering GPT-5.6, with three 301s.

The detector record across the runs:

Confirmed pairs from detector:  0 (every run)
Similarity candidates per run:  13–35
Of 32 candidates on 09-17:      21 were per-tool pricing pages (same title template, different intent)
Merges shipped:                 21, each confirmed by manual triage

The shared-query gate never fires on a site this size. Fresh duplicates are under a month old and have almost no query data yet. What finds them is grouping articles by model family ("GPT-5.x", "Gemini 3.x", "Claude …") and asking one question per pair: same intent, same tool set? The "how much does Claude cost" topic regrew a duplicate three times. It now gets its own check on every run.

A second manual method worked once the detector gave up: a Search Console split on query × page, keeping queries where more than one URL has 2+ impressions. That found a pair where Google had already picked a winner (position 5.9 vs 11.3).

The cadence I ended up with

None of this was designed upfront. It grew from failures.

# content maintenance loop, as it runs today
performance_check:
  when: daily 06:30
  target: articles published ~7 days ago
  output: grade A–D → refresh queue (refresh | title_meta_fix)

refresh_run:
  when: every 3–4 days            # 14 runs in 54 days
  cap: 3–5 refreshes + ≤3 merges
  sources:
    - refresh queue (from the 7-day check)
    - stale sweep: pages with impressions, oldest facts first (85–96 candidates waiting)
    - model-family grouping for merge candidates
  skip: anything refreshed in the last ~11 days
  after: IndexNow submit, verify live render

The queue emptied repeatedly: 22 entries processed, 0 pending at the latest run. The stale sweep never empties. It has held 85–96 candidates the whole time. At 4–5 refreshes per run, twice a week, the backlog takes about 10 weeks to clear. That is longer than the median lifespan of the facts inside it.

A side effect: refreshes make articles longer. 970→1690, 1254→1510, 1363→1714 words. Each refresh adds a "now X, previously Y" paragraph, and those paragraphs will rot too.

What I'd change: facts as data, prose as a view

The root problem is that prices and version numbers live inside prose, copied into dozens of articles. One Copilot price change means grepping 300 bodies and hand-editing each one. Merges exposed this constantly: the "winner" article often held stale data that the "loser" had already fixed.

The structure I'm moving to:

1. Evergreen hubs plus short-lived version posts. "Claude vs ChatGPT for coding" is the hub. It never names a specific version in its title. "GPT-5.6 Terra: what changed" is a version post with an expected life of about 60 days. When the next version ships, it gets merged into the hub and 301'd. That is the same thing I was already doing by hand, now planned from the start.

2. A fact registry. Every volatile claim lives in one place, with a source and a verified date:

# facts/pricing.yaml
copilot_pro:
  value: { annual_usd_month: 10, monthly_usd_month: 15 }   # two billing modes, store both
  billing: usage-based AI credits since 2026-06
  source: https://github.com/features/copilot/plans
  verified: 2026-09-14
  ttl_days: 30

gemini_code_assist_free:
  value: null                      # tier removed
  ended: 2026-06-18
  replacement: "Standard ~$19–23/user via Google Cloud"
  verified: 2026-09-17
  ttl_days: 90

sonnet_5_api:
  value: { input_per_mtok: 2, output_per_mtok: 10 }
  note: intro pricing made permanent 2026-08-10
  verified: 2026-09-24
  ttl_days: 30

Articles reference it instead of hardcoding values:

Copilot Pro costs {{ fact "copilot_pro.annual_usd_month" }} USD/month on annual billing
<small>(verified {{ fact_date "copilot_pro" }})</small>.

The detection problem becomes a query:

# stale facts drive the refresh queue, not article age
stale = [k for k, f in registry.items()
         if (today - f["verified"]).days > f["ttl_days"]]
pages = [p for p in pages if p.facts_used & set(stale)]

One price change, one file edit, every page updated on the next render. The refresh run turns into "re-verify 12 expired facts" instead of "reread 96 articles to find the wrong sentence".

3. Date-stamp every volatile claim in the rendered text. "As of 2026-09" costs four words. It turns a false statement into an old one, and readers forgive old much faster than false.

The 46-day median comes from how long it takes a human-shaped loop to reread prose. A TTL on a fact expires on schedule, whether or not anyone rereads the page.