← Back to lab

598 Dead API Calls a Day: My Retry Cap Only Counted Successes

A drained API key exposed three retry bugs: a --max flag that ignored failures, 1,400 rows stuck unreachable in PENDING, and a circuit breaker covering one consumer.

598 Dead API Calls a Day: My Retry Cap Only Counted Successes
Systems: 2 (a scheduled generator, an enrichment worker)  |  Bugs: 3  |  Window: 2026-08 to 2026-09  |  Status: loops fixed, zombie sweep still open

Two scheduled generation jobs made roughly 598 model API calls a day, every day, and every single one returned 403. The jobs had a --max flag capping how many calls they should make. It worked as written — it just didn't count failures. Once the key died, the only stop condition in the loop became unreachable.

I found it on 2026-09-30 while tracing why a shared OpenRouter key showed $0 of $30 remaining with no reset. The flag was --max=20-shaped, the loop was made += 1-shaped, and made only incremented on a 200. A 403 was free. Free means unbounded.

The cron that couldn't stop

Stripped of everything project-specific, the loop was:

made = 0
for item in queue:
    r = model(item)
    if r.ok:
        made += 1          # 403 never touches this
    if made >= MAX_CALLS:  # never true again once the key is dead
        return

Two jobs, scheduled at 13:00 and 13:30, walking a prompt queue and a glossary queue. The key behind them was shared by three projects and had burned through its total limit. Every call after that returned 403 Key limit exceeded (total limit) — an error no amount of retrying can ever fix, because the failure isn't transient, it's terminal.

That's the part that makes this bug class expensive rather than merely wasteful: the loop wasn't retrying in hope of recovery, it was re-attempting a permanently failed operation, once per queue item per day, forever. ~598 requests a day producing nothing. And because each call "failed cleanly," no alert fired — the jobs exited 0.

The fix was two lines. Count attempts, not successes:

made += 1  # every call costs a call, 200 or 403

…plus removing the retired models from the scheduler config entirely. But the deeper lesson came when I went looking for the same pattern elsewhere and found it wearing a different costume.

The rows that couldn't die

A separate project — an e-commerce enrichment worker — had been effectively idle since 2026-08-03 — a 403 key-limit plus a 402 from the fallback provider killed its pipeline. The worker's own logs said processedCount: 0 every five minutes, which reads as "no work available."

The table had thousands of PENDING rows.

The claim query:

SELECT ... WHERE status = 'PENDING' AND retry_count < 3;

And the recovery job that runs when a PROCESSING row looks stale (crashed worker, killed process):

UPDATE rows
SET status = 'PENDING',
    retry_count = retry_count + 1   -- unconditional
WHERE status = 'PROCESSING' AND stale_since > X;

Follow one row. It's PROCESSING at retry_count = 2 when the worker dies. Recovery moves it back to PENDING and bumps it to 3. The claim query wants retry_count < 3. 3 < 3 is false. The row is PENDING, which the dashboard renders as "queued," and no query will ever select it again. The only code path that marks a row FAILED fires while the row is still PROCESSING — a state it has already left.

Terminal state: unreachable. 2,227 rows were parked there as of 2026-08-16. Five weeks later, 1,401 were still there. Nothing fails them, nothing processes them, nothing reports them. They're invisible from every angle except a hand-written SQL count.

Same bug, different representation: the cron loop had a ceiling that failures didn't count toward, so it ran unbounded. The worker had a ceiling that a recovery path could push rows past, so they never terminated. A retry limit that doesn't bound attempts isn't a limit — it's a comment.

There was a second feeder into the same black hole, found 2026-09-27: a fast-path in the writer that commits its results and returns without marking the row PROCESSED. The row looks stale, recovery re-queues it, the next worker run re-enriches it — up to three extra paid API calls per row — and then parks it as a zombie at retry_count = 3. Recovery code that assumes "stale means no side effects happened" will happily re-run work that did.

The breaker that covered one consumer

The stack even had a quota guard — a circuit breaker that trips on repeated 4xx from the model API and pauses claiming. It worked. It paused raw-import claiming.

The SEO-description queue called the same dead key through its own path and kept hammering it. The breaker lived in one consumer instead of in the HTTP client, so it protected one queue and missed every other caller. A circuit breaker per consumer is a fuse on one wire of a parallel circuit.

The amplifier

One more number from the same audit, because it shows how retry accounting compounds. The enrichment worker processed products in batches of 10 with an all-or-nothing commit: one bad product threw, and all ten rows reset for retry. Result: 97% of the table — about 97,000 rows — sat in FAILED, growing 3,000 to 12,000 rows a day, with no error column ever populated. The failure rate wasn't 97% because 97% of products were bad. It was one-in-ten bad rows poisoning their neighbors, re-tried as a group, failing as a group.

SystemSymptomRoot causeScale
Scheduled generator~598 calls/day, all 403--max counted successes onlydrained a $30 key
Enrichment workerprocessedCount: 0 with a full queuerecovery pushes retry_count past the claim ceiling1,401–2,227 zombie rows
Batch writer97% FAILED, growing dailyone bad item resets the batch of 10~97k rows

Enrichment progress sat flat at 50.1% the whole time, which is the number every dashboard showed while none of this was visible.

What I'd change

The two invariants I now write down before writing retry logic: every attempt increments the counter, whatever the response code; and every reachable non-terminal state must have an edge to a terminal one — if PENDING requires retry_count < max, then no writer may produce PENDING with retry_count >= max, and a sweep should enforce it:

UPDATE rows SET status = 'FAILED', error = 'unreachable-sweep'
WHERE status = 'PENDING' AND retry_count >= 3;

Beyond that: the breaker belongs in the shared HTTP client where it covers every consumer, batches need per-item commits, and if the failure isn't persisted next to the row, the outage is undiagnosable from the database — which is exactly how 97,000 silent failures accumulated before anyone looked.