← Back to lab

Four Env Vars I Never Set Broke My Agents. Three Failed Silently.

PORT, NODE_ENV, ALL_PROXY and leaked CLAUDE_CODE_* — inherited shell env took down an API, broke npm installs and hung a fetch forever. Env is an implicit input.

Four Env Vars I Never Set Broke My Agents. Three Failed Silently.
Env vars: 4 (PORT, NODE_ENV, ALL_PROXY, CLAUDE_CODE_*)  |  Incidents: 5  |  Window: 2026-09-28 → 2026-10-05  |  Status: all fixed

pm2 restart api --update-env took an unrelated API down. The shell that ran it exports PORT=3002 globally — the port of a UI server living on the same box, set in the login profile years ago. --update-env re-reads the shell environment, injects that PORT, dotenv refuses to override an already-set variable, the process fights the UI server for the port, loses with EADDRINUSE, and lands in PM2 errored. The API's own config said port 3006 the whole time. The code was correct; the environment wasn't.

That was the fourth inherited environment variable in a week to break agent work on this server. I logged them all. One failed loudly. The rest didn't fail at all — they just produced wrong results quietly.

The setup that makes this inevitable: agents share a login shell and a process manager with production ops. Every export intended for one service becomes an implicit argument to every command every agent ever runs. Nobody reviews these inputs because nobody "set them in the code."

1. PORT — the restart that took the API down

# login profile: PORT=3002 for the UI server. Global. For everything.
pm2 restart api --update-env
# → injects PORT=3002 → dotenv won't override → EADDRINUSE → errored

PORT=3006 pm2 restart api --update-env   # the fix, every single time
curl -s localhost:3006/health            # never skip the verify

The trap is the word "update" in --update-env. It reads as refresh from the .env file; it actually means re-snapshot whatever shell you're standing in. Any stray export in a profile becomes production config for a service that never asked for it.

2. NODE_ENV=production — two failure modes, neither obvious

The agent shell runs with NODE_ENV=production. That's correct for the servers it hosts and poison for everything it builds.

Mode one: npm quietly drops your devDependencies. Building a CLI from source for a hands-on article test, npm ci completed clean — then the build died with exit 127, cross-env: not found. The bin was never installed. Under NODE_ENV=production, npm treats devDependencies as intentionally omitted:

npm config get omit
# dev          ← npm "working as intended", zero warnings
unset NODE_ENV && npm ci --include=dev   # the fix

Install exits 0. The failure surfaces two steps later as a missing binary, pointing you at the build system instead of the environment. That one cost at least two failed background build attempts before diagnosis.

Mode two: tests fail with a React error on a repo you didn't break. Ambient NODE_ENV=production flips Vite's export-condition resolution to React's production bundle, and Testing Library dies with act(...) is not supported in production builds of React. There's a bonus stage: any later npm install <pkg> under production prunes the already-installed devDependencies — vitest vanishes from node_modules/.bin, and the next npm test reports vitest: not found on a working checkout.

NODE_ENV=development npx vitest run    # pin it per invocation

3. ALL_PROXY=socks5h — the fetch that never returned

A worker running under PM2 fetched a daily exchange rate from a central-bank endpoint. The process environment carried ALL_PROXY=socks5h. The HTTP client honored it. The request didn't fail — it hung, indefinitely, because the call had no timeout. No log line, no error, no retry counter incrementing. The pipeline simply stopped progressing while looking perfectly healthy from the outside.

// every outbound call, explicitly:
const res = await axios.get(url, { proxy: false, timeout: 10_000 });

Silent-failure budget spent: 100%. An unconfigured proxy var plus a missing timeout is the closest thing to an invisible bug a networked worker can have.

4. Leaked CLAUDE_CODE_* — restarts never forget

An agent shell had leaked CLAUDE_CODE_* variables into a PM2 process's environment. Here's the part worth knowing: `pm2 restart --update-env` only adds and updates keys. It never removes them. The stale vars survived every restart, and no amount of .env hygiene on disk changed what the process actually saw.

pm2 delete app && pm2 start ecosystem.config.js && pm2 save
# delete + fresh start is the only reliable env reset

Leaked keys behave like state you can't see in any config file — they exist only in PM2's saved process snapshot.

5. The diagnostic that worked: env -i

A Google Indexing API push script had a 403 habit — intermittent, never explainable by permissions or quotas. Running it under a scrubbed environment, with 3-second spacing between submits, went 5/5 accepted. Zero 403s. Same script, same key, same day. The only variable that changed was the environment it inherited.

env -i HOME="$HOME" PATH="$PATH" bash push-script.sh

That's the inversion worth internalizing: the inherited environment isn't just a source of breakage, it's the first suspect when behavior makes no sense. Scrubbing it is a five-second experiment that either fixes the thing or eliminates half the hypothesis space.

What actually broke

VariableSet byDamageFailure mode
PORT=3002login profile (UI server)API restart → EADDRINUSE → erroredloud
NODE_ENV=productionagent shellnpm ci omits devDeps → exit 127; act() dies; devDeps prunedsilent, delayed
ALL_PROXY=socks5hPM2 process envFX fetch hangs forever, no timeoutsilent, total
CLAUDE_CODE_*leaked from an agent shellstale vars persist across every restartsilent

What I'd change

Treat the environment as a function signature — explicit arguments, no ambient inheritance:

  • Critical scripts run under env -i with an explicit allowlist. Case 5 proved the diagnostic value; now it's the default for anything that talks to an API with auth attached.
  • Restarts always carry explicit VAR=x pm2 restart --update-env. Bare --update-env from an interactive shell is banned.
  • Every outbound fetch pins timeout and proxy behavior per call. A request without a timeout is a future zombie process.
  • npm and test-runner invocations pin NODE_ENV per command.
  • Before debugging any "impossible" failure on a shared box: env | sort and read it. The bug is in the code about as often as you'd expect.

The uncomfortable summary: four production incidents in eight days, zero caused by code. Every one was an input nobody remembers setting.