← Back to lab

Silent Success Is the Worst Bug: 8 Tools That Lied to My Agents

Exit code 0, HTTP 200, 'Uploaded.' Eight real cases where my agents got a success signal for work that never happened, and the guard I added for each one.

Silent Success Is the Worst Bug: 8 Tools That Lied to My Agents

Exit 0 Is Not a Result

Eight times in the last five months an agent of mine reported "done" on work that was broken. Each time the agent read the signal correctly: exit code 0, HTTP 200, a green "success" line. The signal was wrong.

My agents run unattended: scheduled content refreshes, tool-catalog updates, uploads, deploys. Nobody watches the terminal. A loud failure gets caught at the next check. A quiet failure ships, gets logged as success, and turns into a precedent for the next run.

Model: Claude (Claude Code, headless + interactive)  |  Period: 2026-04 to 2026-09  |  Incidents: 8  |  Status: guarded

Every incident below has the same four parts: what the agent saw, what was actually true, how far it spread, and the guard.

1. wrangler r2 object put without --remote

Agent saw: a success message for the upload and a public URL built from the bucket base plus the key.

Truth: without --remote, wrangler writes to a local simulation of the bucket. It still prints success. The object never reaches Cloudflare. Every link 404s.

Blast radius: every link handed to me from that upload went nowhere. The agent had no reason to doubt them, because the tool told it the upload worked.

Guard: a PreToolUse hook blocks the command, and all uploads go through one wrapper that prints a URL only after it returns 200.

# PreToolUse hook (Bash matcher), simplified
if echo "$CMD" | grep -q 'wrangler r2 object put' && ! echo "$CMD" | grep -q -- '--remote'; then
  echo "Blocked: without --remote this writes to a local simulation. Use r2-upload.sh." >&2
  exit 2
fi
# wrapper: upload, then prove it
url="$PUBLIC_BASE/$key"
code=$(curl -s -o /dev/null -w '%{http_code}' "$url")
[ "$code" = "200" ] && echo "$url" || { echo "UPLOAD NOT VISIBLE ($code)" >&2; exit 1; }

2. A collection API that returns 20 of 55

Agent saw: a 200 from a content API I built, with a clean JSON list of AI tools. No error, no truncation marker.

Truth: the endpoint returned 20 items. The table held 55 active ones. No totalItems, no next link, and raising itemsPerPage returned the same 20. 35 records (64%) were invisible to anything that read the API.

Blast radius: the catalog skill checked for duplicates against the API before creating a tool. It could not see existing entries, so it created them again. Result: 8 empty duplicate stubs, including one tool in 3 copies and another in 3 copies. They showed up in the public listing and made the article-to-tool linker attach the same article to several rows.

Guard: inventory and dedup read from the database, not the API. On a sister endpoint that does paginate (capped at 20 per page no matter what you ask for), the loop runs until next disappears. Any list call without a total count is treated as partial.

# never trust a page as the whole set
items, page = [], 1
while True:
    r = get(f"/api/content_items?page={page}")
    items += r["member"]            # not "hydra:member", despite the docs
    if "next" not in r.get("view", {}):
        break
    page += 1

3. 302 -> /login counted as "page works"

Agent saw: the new admin route was registered in the router, and a curl returned 302 to the login page. It reported "route works, just needs login."

Truth: a 302 to login proves the firewall caught a registered route. That is all. The first real authenticated hit returned 500: a TypeError in the template's favicon block, because the page extended an admin framework layout whose context only exists inside that framework's CRUD routing.

Blast radius: I was the first one to click it, and I got the 500.

Guard: for protected pages, "verified" means an authenticated render (curl with a session cookie) or a template lint plus a tail of the prod log right after a real hit. Route listing and redirect codes count for nothing.

4. A cached prod env holding a revoked key

Agent saw: a valid OPENROUTER_API_KEY in the app's .env.local, tested against the API, 200. The generation command still failed with 401.

Truth: the framework compiles env files into a cached PHP file in production. That cache held an older, revoked key. Running .env.local through a manual check proved nothing about the value the app was actually using.

Blast radius: a batch generation command that failed on every call. Clearing the prod cache on a live site was rejected as too disruptive.

Guard: real environment variables override the compiled cache, so the fix is an inline override. The rule: test the key the process actually reads, not the key in the file you are looking at.

KEY=$(grep -E '^OPENROUTER_API_KEY=' .env.local | cut -d= -f2- | tr -d "\"'")
OPENROUTER_API_KEY="$KEY" php bin/console app:pseo:generate --tool="$SLUG" --force

5. Literal \n stored as the article body

Agent saw: during a content refresh, an existing article that looked fine in the API: 200, 1,254 words, correct title.

Truth: the body contained literal \n escape sequences instead of newlines. The markdown renderer got one giant paragraph. Zero headings rendered on the live page.

Blast radius: a published article with no structure, sitting live until a refresh happened to read it. Whichever write introduced it got a 200 back. The API stored the string it was given.

Guard: a pre-write check on anything headed for a markdown field, plus a post-publish count of rendered <h2> tags.

assert "\\n" not in body, "literal backslash-n in body: double-escaped upstream"
assert body.count("\n## ") >= 3, "suspiciously few headings"

6. A refresh that saved but never showed

Agent saw: PATCH on the article body, HTTP 200, the new text in the DB row, publish date bumped. Everything an API can say to confirm a write.

Truth: tutorial-type items render from a structured contentStructure field. The markdown body is only a fallback, used when that field is empty. This article had a populated but broken structure: two throwaway sections, one of them using the wrong keys, so it rendered nothing. The 1,118 → 1,537 word refresh was stored and never shown. The live page stayed a stub.

Blast radius: a "refreshed" article that nobody could read. On a second pass I hit a variant: email-gated items render only a preview to anonymous visitors, so grepping the public page for new text fails even when the update worked.

Guard: before refreshing, check type and contentStructure. After refreshing, curl the public URL and grep for a phrase that only exists in the new version. For gated items, check the API plus a gate marker on the page instead.

7. curl ran, then the script crashed

Agent saw: a Python script that created a catalog entry crashed with a traceback. It fixed the parsing bug and ran it again.

Truth: the crash happened after subprocess.run(["curl", ... "-X", "POST" ...]). The POST had already landed. The retry created the entry a second time, and the API quietly gave it a -1 slug suffix.

Blast radius: one duplicate public entry, found in the same run and deleted (DELETE, 204). Caught early only because the -1 looked wrong.

Guard: after a crash in any script that made a write, check for the record before retrying. And treat a numeric slug suffix on create as a duplicate alarm, not a naming detail.

created = post_tool(payload)
if re.search(r"-\d+$", created["slug"]):
    raise RuntimeError(f"slug {created['slug']} looks like a duplicate, check before continuing")

8. A token-saving CLI proxy truncated stdout

Agent saw: API responses from curl -s that were valid-looking but short. JSON parsing failed, or worse, parsed a partial object.

Truth: a CLI proxy I run to cut token usage started truncating curl stdout to about 583 bytes. The full body went to a tee log the agent didn't know existed. Nothing printed a warning.

Blast radius: several wasted calls per run and some false alarms: "field missing" when the field was past the cut.

Guard: API responses go to a file. The agent parses the file, never stdout. Same fix for the Telegram bot API, where piping unicode-heavy responses into json.load(sys.stdin) choked.

curl -s -o /tmp/resp.json -w '%{http_code}\n' "$API/content_items/$ID"
python3 -c 'import json; d=json.load(open("/tmp/resp.json")); print(len(d["fullContent"].split()))'

Where the Pattern Holds

#Signal the agent trustedLayer that liedDetected by
1"Uploaded"CLI default mode404 on the link
2200 + listAPI serializerDB count mismatch
3302Auth firewallHuman click, 500
4Key in .env.local worksCompiled env cache401 at runtime
5200 on readUpstream escapingRendered heading count
6200 on PATCHTemplate branchPublic page grep
7Traceback = nothing happenedSide effect before crash-1 slug
8curl stdoutOutput proxyParse failure

All eight passed the check the agent actually ran. One was caught by a human clicking, one by chance during unrelated work.

My original prompts had rules like "check the response status". The agent followed them every time, and the status was fine every time. What I'd change: I wrote the verification rules after each incident, one at a time. They should have been one rule from the start, applied to every write:

- Verify at the destination, not the tool output.
  Upload    -> GET the public URL, expect 200
  DB write  -> read it back through the path users see
  Page      -> authenticated render + log tail, not the route table
  Secret    -> test the value the running process reads
  List      -> compare against a count from the source of truth
- A crash after a write is not a no-op. Look before you retry.
- Parse files, not stdout.