Products: 311,447 | Backlog: 8,074 → 2,074 | LLM-mapped pairs: 2,380 | Build: 2 days | Date: 2026-10 | Status: deployed A category-mapping script on my price-aggregation backend exited green every time it ran. The backlog it was supposed to clear — 8,074 products carrying a vendor category string but NULL for the internal category — never moved. The script worked. Its output just didn't survive the next feed ingest.
The catalog holds 311,447 products from 14 active vendor feeds (209 vendors, several arriving through aggregators). Vendors send free-text category chains like Pet > Dogs > Dry food; the internal taxonomy is what facets, filters, and category pages read. Products that land in no internal category are invisible to every browse path — they only surface in search. 8,074 of them sat there through repeated "successful" mapping runs.
Why nothing stuck
The old flow had two writers and one reader. A one-off script mapped raw vendor categories to internal IDs in its own table. Meanwhile, every feed ingest re-assigned categories through a different path that reads a mapped_category_id column on the vendor-category table — a column the script never touched. Each ingest re-derived what it could and left the rest NULL. The script's mappings were correct and orphaned.
The fix wasn't a better mapper. It was writing to the place ingest actually reads:
-- confirmed mappings propagate to the column the ingest writer reads,
-- only where NULL — a manual decision is never overwritten by a machine one
UPDATE vendor_categories
SET mapped_category_id = m.category_id
FROM category_mappings m
WHERE vendor_categories.mapped_category_id IS NULL
AND m.category_id IS NOT NULL; After that, a mapping survives every feed run, and new vendor categories flow through automatically.
The pipeline
| Stage | What it does | Cost control |
| Tree build | Existing 136 nodes as skeleton; high-demand unmapped vendor categories added as leaves, LLM picks the parent | LLM capped, idempotent upsert |
| Demand weighting | search_interest backfilled per category from my shared SERP proxy | hard cap 200 queries per run |
| Exact match | Normalized name/slug equality → confidence 1.0, method exact | free |
| LLM bulk | Remaining pairs batched to a cheap reasoning model, strict JSON out | ≤80 calls per run |
| Gate | Confidence ≥ 0.7 confirmed; everything else → review queue | — |
| Weekly job | Re-runs the whole thing idempotently for new feed categories | env kill-switch |
The tree grew from 136 to 411 nodes. 200 categories carry a search-interest score; the cap means the other 211 are structurally present but unweighted until a later run.
The model truncates its own output
The interesting failure was in the LLM stage. Batches of 100 pairs, temperature: 0, max_tokens: 16000, strict JSON array promised. The model (deepseek-flash, the cheap end of the pool) is a reasoning variant — it emits reasoning tokens before the JSON. Reasoning eats the budget. The array arrives cut mid-object:
[
{"index": 1, "category_path": "Pet > Dogs", "confidence": 0.9},
{"index": 2, "category_path": "Electronics", "confidence": 0.8},
{"index": 3, "categ JSON.parse fails, and here is the part that cost me a run to understand: the generic chatJson helper's built-in repair cannot fix this. JSON repair handles malformed output — markdown fences, prose wrappers, trailing commas. Truncation isn't malformed, it's missing. There is nothing to repair "[{a}, {b}, {c" into; the only honest move is to close what's complete and admit the rest never arrived.
Two fixes, both shipped:
// 50 (not 100): the model reasons before emitting JSON and truncates
// longer arrays at max_tokens — smaller batches plus the repair below
// keep every batch parseable.
const LLM_BATCH_SIZE = 50; // JSON.parse failed on the raw slice — the model ran out of tokens
// mid-array. Salvage the complete prefix instead of dropping the batch.
const start = text.indexOf('[');
const lastClose = text.lastIndexOf('}');
if (lastClose > start) {
arr = JSON.parse(`${text.slice(start, lastClose + 1)}]`);
} The prefix items are valid and kept. The pairs that never came back are treated as never attempted — which matters for the next rule.
Trust boundaries, written down
Three gates, each one learned the hard way elsewhere:
Hallucinated paths are downgraded, not trusted. The model returns category_path strings. The parser validates every one against the actual tree — a path that doesn't exist becomes unmatched, confidence gets clamped to [0, 1], out-of-range indices are rejected. The model never gets to invent taxonomy.
Coverage lies are impossible. Pairs below 0.7 confidence, unmatched pairs, and never-attempted pairs (call cap hit, provider down, output truncated past repair) all land in the review queue as explicit rows. A capped or half-failed run cannot present itself as "nothing left to map" — the missing work is countable.
Ingest owns the final word, but not over humans. Mappings propagate only into NULL slots. A manual confirmation in the review queue is terminal.
Results
From the live run:
| Metric | Value |
| Backlog (path set, category NULL) | 8,074 → 2,074 (−74%) |
| Products with an internal category | 265,172 / 311,447 (85.1%) |
| Vendor categories mapped | 2,578 / 3,061 (84.2%) |
| Confirmed mapping pairs | 2,721 — 341 exact + 2,380 LLM |
| Average LLM confidence | 0.83 |
| Review queue | 1,108 raw categories |
The exact-match fast path mapped 341 pairs for zero cost — 12.5% of all confirmed mappings, free. The remaining 2,074 backlog products trace to raw categories sitting in the review queue.
What I'd change
The review queue is 1,108 raw categories — 29% of all pairs routed to a human. That's too high. A second LLM pass fed with already-resolved sibling categories from the same vendor (few-shot, not zero-shot) would cut it; vendors are internally consistent, and the first pass throws that consistency away. I'd also measure token consumption on the first batch before committing to a batch size — 100 was optimism, 50 is measurement. And the demand-weighting cap leaves 211 of 411 categories unweighted, which means the tree's growth is still steered by whatever happened to be in the feeds rather than by search demand.