← Back to lab

8,074 Products, One Frozen Backlog, and an LLM That Truncates Its Own JSON

A category backlog of 8,074 products never moved while the mapping script ran green. Its output didn't survive ingest, and the LLM truncated its own JSON.

8,074 Products, One Frozen Backlog, and an LLM That Truncates Its Own JSON
Products: 311,447  |  Backlog: 8,074 → 2,074  |  LLM-mapped pairs: 2,380  |  Build: 2 days  |  Date: 2026-10  |  Status: deployed

A category-mapping script on my price-aggregation backend exited green every time it ran. The backlog it was supposed to clear — 8,074 products carrying a vendor category string but NULL for the internal category — never moved. The script worked. Its output just didn't survive the next feed ingest.

The catalog holds 311,447 products from 14 active vendor feeds (209 vendors, several arriving through aggregators). Vendors send free-text category chains like Pet > Dogs > Dry food; the internal taxonomy is what facets, filters, and category pages read. Products that land in no internal category are invisible to every browse path — they only surface in search. 8,074 of them sat there through repeated "successful" mapping runs.

Why nothing stuck

The old flow had two writers and one reader. A one-off script mapped raw vendor categories to internal IDs in its own table. Meanwhile, every feed ingest re-assigned categories through a different path that reads a mapped_category_id column on the vendor-category table — a column the script never touched. Each ingest re-derived what it could and left the rest NULL. The script's mappings were correct and orphaned.

The fix wasn't a better mapper. It was writing to the place ingest actually reads:

-- confirmed mappings propagate to the column the ingest writer reads,
-- only where NULL — a manual decision is never overwritten by a machine one
UPDATE vendor_categories
SET mapped_category_id = m.category_id
FROM category_mappings m
WHERE vendor_categories.mapped_category_id IS NULL
  AND m.category_id IS NOT NULL;

After that, a mapping survives every feed run, and new vendor categories flow through automatically.

The pipeline

StageWhat it doesCost control
Tree buildExisting 136 nodes as skeleton; high-demand unmapped vendor categories added as leaves, LLM picks the parentLLM capped, idempotent upsert
Demand weightingsearch_interest backfilled per category from my shared SERP proxyhard cap 200 queries per run
Exact matchNormalized name/slug equality → confidence 1.0, method exactfree
LLM bulkRemaining pairs batched to a cheap reasoning model, strict JSON out≤80 calls per run
GateConfidence ≥ 0.7 confirmed; everything else → review queue—
Weekly jobRe-runs the whole thing idempotently for new feed categoriesenv kill-switch

The tree grew from 136 to 411 nodes. 200 categories carry a search-interest score; the cap means the other 211 are structurally present but unweighted until a later run.

The model truncates its own output

The interesting failure was in the LLM stage. Batches of 100 pairs, temperature: 0, max_tokens: 16000, strict JSON array promised. The model (deepseek-flash, the cheap end of the pool) is a reasoning variant — it emits reasoning tokens before the JSON. Reasoning eats the budget. The array arrives cut mid-object:

[
  {"index": 1, "category_path": "Pet > Dogs", "confidence": 0.9},
  {"index": 2, "category_path": "Electronics", "confidence": 0.8},
  {"index": 3, "categ

JSON.parse fails, and here is the part that cost me a run to understand: the generic chatJson helper's built-in repair cannot fix this. JSON repair handles malformed output — markdown fences, prose wrappers, trailing commas. Truncation isn't malformed, it's missing. There is nothing to repair "[{a}, {b}, {c" into; the only honest move is to close what's complete and admit the rest never arrived.

Two fixes, both shipped:

// 50 (not 100): the model reasons before emitting JSON and truncates
// longer arrays at max_tokens — smaller batches plus the repair below
// keep every batch parseable.
const LLM_BATCH_SIZE = 50;
// JSON.parse failed on the raw slice — the model ran out of tokens
// mid-array. Salvage the complete prefix instead of dropping the batch.
const start = text.indexOf('[');
const lastClose = text.lastIndexOf('}');
if (lastClose > start) {
  arr = JSON.parse(`${text.slice(start, lastClose + 1)}]`);
}

The prefix items are valid and kept. The pairs that never came back are treated as never attempted — which matters for the next rule.

Trust boundaries, written down

Three gates, each one learned the hard way elsewhere:

Hallucinated paths are downgraded, not trusted. The model returns category_path strings. The parser validates every one against the actual tree — a path that doesn't exist becomes unmatched, confidence gets clamped to [0, 1], out-of-range indices are rejected. The model never gets to invent taxonomy.

Coverage lies are impossible. Pairs below 0.7 confidence, unmatched pairs, and never-attempted pairs (call cap hit, provider down, output truncated past repair) all land in the review queue as explicit rows. A capped or half-failed run cannot present itself as "nothing left to map" — the missing work is countable.

Ingest owns the final word, but not over humans. Mappings propagate only into NULL slots. A manual confirmation in the review queue is terminal.

Results

From the live run:

MetricValue
Backlog (path set, category NULL)8,074 → 2,074 (−74%)
Products with an internal category265,172 / 311,447 (85.1%)
Vendor categories mapped2,578 / 3,061 (84.2%)
Confirmed mapping pairs2,721 — 341 exact + 2,380 LLM
Average LLM confidence0.83
Review queue1,108 raw categories

The exact-match fast path mapped 341 pairs for zero cost — 12.5% of all confirmed mappings, free. The remaining 2,074 backlog products trace to raw categories sitting in the review queue.

What I'd change

The review queue is 1,108 raw categories — 29% of all pairs routed to a human. That's too high. A second LLM pass fed with already-resolved sibling categories from the same vendor (few-shot, not zero-shot) would cut it; vendors are internally consistent, and the first pass throws that consistency away. I'd also measure token consumption on the first batch before committing to a batch size — 100 was optimism, 50 is measurement. And the demand-weighting cap leaves 211 of 411 categories unweighted, which means the tree's growth is still steered by whatever happened to be in the feeds rather than by search demand.