185,342 active offers in the database. 67,275 offers in the actual vendor feeds. The 118,067 in between — 64% of the catalog — were ghosts, and a single UPDATE statement resurrected them every night.
I run a product aggregator that feeds Czech price-comparison sites. Vendor feeds re-download on a schedule, and a feed has no delete event — only presence and absence. Absence-to-expiry was supposed to be the job of a last_seen timestamp: offers the feed stops carrying go stale, the expiry job deactivates them. Simple, standard, wrong in one detail.
The bug: freshness at the wrong scope
After every successful run, the ingestor refreshed the vendor's rows like this:
// ran on every successful feed download, for the whole vendor
await db.query(
`UPDATE offers SET last_seen = now()
WHERE vendor_id IN ($1, $2)` // uuid key OR legacy slug key
); Presence in the feed was never part of the predicate. Any successful download touched every offer that vendor had ever produced. A shop that delisted 40% of its range kept all of those offers "fresh" indefinitely, because the shrunken feed still downloaded fine.
The expiry job — deactivate offers with last_seen older than N days — was structurally incapable of firing. The mechanism designed to detect absence was being reset by nothing more than the vendor still existing.
A legacy dual-key made it worse: old rows keyed by vendor slug matched the same blanket UPDATE, so even offers from a half-finished migration years back got the nightly bump.
The monitor agreed with the bug
/health computed freshness as MAX(last_seen) against a 48-hour threshold.
Measured on July 19: /health reported offers: fresh while 106,531 of 171,454 active offers (62%) had last_seen older than 7 days. The Telegram monitor wired to this signal was alerting on a value that one live offer out of 171k could satisfy. The stale-price alert built specifically to catch this was structurally incapable of firing.
MAX() is a liveness check. Freshness is a distribution. Any dashboard that collapses a per-row property into a max over 171k rows will lie to you exactly when you need it.
The fix
Touch only what the run actually saw, and never expire off a run that looks truncated:
// freshness = (vendor, url, current name/brand) seen in THIS payload
await db.query(
`UPDATE offers SET last_seen = now()
WHERE vendor_id = $1 AND url = ANY($2) AND name = ANY($3)`
);
// circuit breaker: a run delivering <50% of the previous count
// is a truncated download, not a mass delisting — skip expiry
if (runCount < 0.5 * prevCount) skipExpiry(); The scope of the touch is the whole fix. Offers absent from the payload now age out on their own; the 50% fuse stops one broken CDN response from deactivating half a vendor.
Active offers went 185,342 → 67,275 — equal to the sum of the feed sizes for the first time since the project launched. 74 days passed between the first measured symptom and the fix.
What the clean data revealed
The ghosts were hiding real problems. Once the catalog matched reality:
| Metric | Before | After |
| Active offers | 185,342 | 67,275 |
| Ghost offers | 118,067 (64%) | 0 |
| Sellable products without category | 75% | 6% |
| Products without brand | 36% | 1% |
| Products sold by 2+ shops | unknown | 0 |
Zero cross-shop overlap. After cleanup, not a single product is sold by more than one shop. The price-comparison feature — the reason this project exists — has no substrate. Every "compare prices" surface was comparing nothing. One foreign-market megastore accounted for 44% of sellable offers on the Czech sites; it's now hidden behind a market flag (29,897 offers reclassified, 10,561 enrichment-queue items pruned).
Wrong vendor attribution. One consolidated feed host serves three different shops, but raw records were keyed by the feed hostname — so all three shops' offers belonged to one vendor. The real shop id sat in a query parameter on every offer URL. Re-attribution by that parameter merged 27,135 duplicate offers.
Corrupted price history. 257k snapshots with mixed currencies in one series were deleted; 789k EUR snapshots were converted at the central bank's daily reference rate (24.465 CZK/EUR that day). Canonical price is CZK, the vendor's own value survives in source_price/source_currency.
Categories without a model. A rule-based CZ/SK regex mapper, leaf-first with a fallback tier, runs at ingest. No-category went 75% → 6%, no-brand 36% → 1%. Zero AI calls in the loop for what AI had been queued to do.
One debugging detour worth keeping: a new outbound HTTP call in the worker hung forever under PM2 — no error, no timeout, just a pending promise. Cause: ALL_PROXY=socks5h in the process environment, which axios honors silently. The fix is per-call proxy: false. An "axios HTTP/2 quirk" documented months earlier in the downloader was almost certainly this same proxy variable. It worked in every shell test and died only under PM2, because the proxy var lives in the PM2 environment.
What I'd change
Expiry on absence needs the truncated-run fuse, but the real lesson is the dashboard: active offers plotted against sum of feed sizes is one line pair, it costs nothing, and it would have exposed 118k ghosts in week one instead of leaving it to a quarterly audit. Every freshness signal since then is a ratio or a percentile — the only aggregate that ever fired correctly was the one a human computed by hand.