Consumers: 13 (sites, agents, cron jobs) | Paid calls this month: 1 of 100 | Build: 1 day | Date: 2026-10 | Status: live My keyword research endpoint has a dataSource variable. Until October 1, every request left the building with it set to ai-estimated. Nobody noticed — not the cron jobs consuming the output, not the agents building landing pages from it, not me. Every search volume and difficulty score was a language model's guess, serialized in the same JSON shape, under the same field names, as measured data.
I run SEO for a dozen small content sites. The keyword flow is standard: an LLM proposes keyword ideas with rough volume/difficulty estimates, then a real SERP provider is supposed to overwrite the estimates with actual numbers. A fail-soft catch guards the second step. The guard worked perfectly — right up until it silently guarded nothing.
The bug
The code read SERPAPI_KEY. The environment defined SERPAPI_API_KEY. The real-data path never fired, and the fail-soft design made that invisible:
let dataSource = 'ai-estimated';
try {
realData = await getKeywordData(keywordList);
dataSource = 'real';
} catch (err) {
console.error('lookup failed, keeping AI estimates:', err);
}
// keywords ship either way — identical shape, identical fields I found it while wiring up a replacement for that entire call path. Not because anything alerted. Because I was reading the env handling line by line.
Fail-soft is the right availability call — a keyword tool that 500s because SerpApi hiccuped is worse than one that returns estimates. The sin is that degraded output was indistinguishable from real output. estimatedVolume: "high" is the same field whether it's measured or invented. The dataSource mark existed, then died in scope. Nothing downstream could see it.
Rule I now hold: a fallback must label its output where consumers actually look — a response-level flag, a warn-level log, a metric. console.error in a container log is none of those.
The waste audit that motivated the rewrite
The typo was the embarrassing part, but the architecture was worse. Every site, agent, and one-off script held its own provider key. Free tiers are per-key: SerpApi gives 100 searches a month, Serper 2,500 credits once. Two sites researching the same topic paid twice, because a cache inside one process helps exactly one process.
The evidence, once I centralized: in the first 50 hours after the cutover, 13 distinct consuming projects hit the shared proxy. One Czech head term — "nejlepší AI nástroje", "best AI tools" — was requested 7 separate times by different consumers. Before the proxy, that was 7 paid calls. After: 1 paid, 6 from cache.
What the proxy actually is
One local service, SQLite behind it, every SERP consumer in the fleet pointed at it:
- Cache keyed on (kind, query, params) — search, related, autocomplete, trends. Same query from any consumer is one paid call, ever, until it expires.
- In-flight dedup — two agents asking the identical query in the same second share one upstream request.
- Stale-if-error — provider dead or quota gone? Serve the cached copy flagged
stale: trueinstead of failing. - Failover — Serper 402s, the router retries on SerpApi, and soft-disables the key.
- Per-project attribution — a
?project=tag on every call, so I can see which of the 13 consumers actually spends. - A cached-only preview endpoint — agents check whether data already exists for zero cost before triggering a paid fetch.
A side effect I didn't plan for: normalization moved to one place. Day one, the proxy caught a provider returning related-searches double-wrapped ({query: {query: "…"}}). Before, each consumer would have hit that bug alone — or shipped the malformed shape and let the next consumer discover it again.
The quota field that lies
Serper responses include a credits field with your remaining balance. Mine said 1. Four calls later it still said 1. The field doesn't track reality, so the proxy doesn't read it.
Quota state is now derived from the only channel providers can't get wrong — the HTTP status. A 402/403/429 soft-disables that key, fails over, and marks the provider degraded; an admin reset re-enables it after a top-up. Self-reported remaining quota is a hint. The error channel is telemetry.
Numbers
First 50 hours after cutover (usage counters reset with the process; quotas persist in SQLite):
| Metric | Value |
| Requests served | 61 |
| Served from cache | 17 (28%) |
| Consuming projects | 13 |
| Top query: asks / paid calls | 7 / 1 |
| Cache reads absorbed since launch | 77 |
| Errors | 0 |
| SerpApi used in October | 1 of 100 monthly free |
| Serper credits used | 46 of 5,000 (two-key pool) |
One paid SerpApi call for the whole month so far, against a fleet that used to fan out per-key.
What I'd change
The dataSource flag should have been a visible field with an alert attached from day one — silent degradation that produces plausible-looking numbers is a fabrication machine — this one stood until I happened to read the env handling line by line. The 28% cache hit rate is honest but low: it's day-two numbers from consumers that just onboarded, and I should push more of them through the cached-only preview before they trigger paid fetches. And the preview endpoint deserves a CLI wrapper, so agents reach for the free check by default instead of remembering it exists.