← Back to lab

One VPS, 28 Vhosts, Zero Cloudflare Bill: The Stack Behind My AI Content Autopilot

28 nginx vhosts, 17 PM2 apps and 16 scheduled Claude jobs on one 8-vCPU VPS. The scheduler, the Cloudflare self-block fix, R2 pitfalls and IndexNow vs GSC.

One VPS, 28 Vhosts, Zero Cloudflare Bill: The Stack Behind My AI Content Autopilot

One VPS, 28 Vhosts, Zero Cloudflare Bill

Host: 8 vCPU (AMD EPYC Genoa)  |  RAM: 15 GiB  |  Disk: 150 GB  |  Uptime: 12 weeks  |  Status: working

28 nginx vhosts. 17 PM2 processes. 18 scheduled Claude jobs that have run 1,464 times without a human in the loop. One VPS, Cloudflare free plan in front, R2 for artifacts. The monthly infra bill is the VPS and nothing else.

The home lab post covered the Mac Mini and the Slack bridge. That was the interactive side. This is the unattended side: the box that writes, refreshes, reindexes and ships content while I am asleep, and the four infrastructure problems that broke it before it worked.

The layout

                    Internet
                       |
              +-----------------+
              |   Cloudflare    |   free plan, proxied DNS,
              |  (own zone)     |   Bot Fight Mode on
              +--------+--------+
                       |
   +-------------------+----------------------------------+
   |  VPS                                                 |
   |                                                      |
   |  nginx :80/:443 ---- 28 vhosts                       |
   |     |-- static / PHP sites                           |
   |     '-- proxy_pass -> PM2 apps (17 procs)            |
   |                                                      |
   |  PM2: ui-server (Node)                               |
   |     '-- in-process scheduler, 30 s tick              |
   |           '-- tasks.json (18 jobs)                   |
   |                 '-- spawns: claude -p "<prompt>"     |
   |                       |-- curl http://127.0.0.1      |
   |                       |     -H "Host: site-a.example"|
   |                       |-- r2-upload.sh --remote      |
   |                       |-- IndexNow ping             |
   |                       '-- Telegram bot API          |
   |                                                      |
   |  ~/.env  (single source of secrets)                  |
   +------------------------------------------------------+
                       |
                 Cloudflare R2  (PDFs, MP4s, images)
ItemSpec / countCost
VPS (Hetzner cloud, AMD)8 vCPU, 15 GiB RAM, 150 GB disk~EUR 30/month (approx. list price)
Cloudflarefree plan, proxied DNS for every site$0
R2artifacts only, well under free tier$0 (zero egress fees)
nginx vhosts28-
PM2 processes17 (14 online)-
systemd services running40-
Scheduled agent jobs18 (16 on a schedule, 2 manual)LLM usage billed separately
LLM spend is the real cost. Hosting is a rounding error next to it.
## Why PM2 per app, not Docker
Every Node/Astro app is one PM2 process behind one nginx vhost. No compose files, no container networking, no image rebuilds.
```bash
# the entire deploy for a typical app
npm run build && pm2 restart my-app && curl -fsS -o /dev/null -w '%{http_code}\n' http://127.0.0.1:4321/
`
One rule is enforced for the agents too: restarts go through PM2 or systemd, whichever owns the process. Never nohup node server.js &. An orphaned node process holding the port is the single most common way an agent "fixes" something and silently breaks the next deploy.
## An in-process scheduler instead of cron
The scheduler is ~130 lines inside the Node UI server. It ticks every 30 seconds and reads a JSON file.
```js
// lib/scheduler.js (simplified)
setInterval(checkScheduledTasks, 30000);
async function checkScheduledTasks() {
const tasks = loadTasks(); // tasks.json
for (const t of tasks) {
if (t.disabledt.status === 'running') continue;
if (Date.now() < Date.parse(t.nextRun)) continue;
runTask(t); // spawns claude -p with t.prompt, t.workingDir
}
}
`
Each task record carries more than a cron line can:
```json
{
"schedule": "daily@06:30",
"workingDir": "~/projects/site-a",
"model": "sonnet",
"lastStatus": "success",
"lastSessionId": "…",
"consecutiveFailures": 0,
"runCount": 212
}
`
Why not cron:
State lives with the job. lastError, lastSessionId and consecutiveFailures sit next to the schedule. Cron gives you an exit code and a mail spool nobody reads.
No overlap. A 25-minute article job that runs long does not get a second copy started on top of it. The running status is the lock.
Same env as the UI. The server sources ~/.env at boot, so scheduled runs see exactly the variables interactive runs see. Cron's empty environment was the source of half the "works when I run it" bugs.
Schedules mix plain cron (0 9 * * 1,3,5) with a daily@HH:MM shorthand. Right now: 0 disabled, 0 in a failure streak.
## The Cloudflare self-block
Every site is proxied through Cloudflare with Bot Fight Mode on. The agent on the VPS calls the sites' own APIs to publish, refresh and fetch content. From Cloudflare's point of view that is a datacenter IP with a curl user agent hitting a WordPress/Symfony API. It gets challenged. My own server, blocked from my own sites.
First attempt: a WAF custom rule with action Skip for the server IP. Did nothing. Bot Fight Mode on the free plan is not part of the WAF phase that skip rules control, so the challenge still fires.
What worked for my zone: a Configuration Rule matching the server's IPv4 and IPv6 and turning the bot/security features off for that source only. Scoped to my own origin's address, in my own zone.
The better fix is to not go through Cloudflare at all. The origin is on the same box. Talk to it directly:
```bash
#!/usr/bin/env bash
# call own site: loopback when running on the origin, public URL elsewhere
ORIGIN_IP="203.0.113.10" # this box's public IP (documentation range here)
site_curl() {
local host="$1" path="$2"; shift 2
if [[ "$(curl -s --max-time 3 ifconfig.me)" == "$ORIGIN_IP" ]]; then
# nginx picks the vhost from Host; Cloudflare never sees the request
curl -sS -H "Host: $host" "http://127.0.0.1$path" "$@"
else
curl -sS "https://$host$path" "$@"
fi
}
site_curl site-a.example /wp-json/wp/v2/posts -H "Authorization: Bearer $SITE_A_TOKEN"
`
Gotchas that each cost an evening:
Sites that force HTTPS in nginx return a 301 on :80. Either exempt 127.0.0.1 from the redirect or use curl --resolve site-a.example:443:127.0.0.1 https://… so SNI and certs still match.
Frameworks that build absolute URLs from the request scheme will emit http:// links in API responses. Set X-Forwarded-Proto: https on the loopback call.
A redirect or a WAF HTML page is not a failure signal by itself. Check the response body and status of the write, not the first hop.
## R2 and the upload that lies
Anything I need to look at on my phone (PDF, MP4, rendered card) goes to R2 and the link goes to Telegram. The trap:
```bash
# writes to a LOCAL simulation. prints success. link 404s.
wrangler r2 object put bucket/key.pdf --file key.pdf
# actually uploads
wrangler r2 object put bucket/key.pdf --file key.pdf --remote
`
Without --remote, wrangler happily writes into its local Miniflare state and reports success. The agent then sends a dead link, confidently. It happened enough times that the bare command is now blocked by a PreToolUse hook, and all uploads go through one wrapper that refuses to print a URL it has not verified:
```bash
up() {
wrangler r2 object put "$BUCKET/$2" --file "$1" --content-type "$(ctype "$1")" --remote >/dev/null \
{ echo "UPLOAD-FAIL $1" >&2; return 1; }
local url="$PUBLIC_BASE/$2"
[[ "$(curl -s -o /dev/null -w '%{http_code}' "$url")" == 200 ]] \
&& echo "$url"{ echo "NOT-200 $url" >&2; return 1; }
}
`
Rule for every agent: never send a URL you have not curl-checked. And never let the model compose the public hostname itself. It will invent a plausible bucket subdomain and it will be wrong.
## One .env, one restart
All secrets live in a single ~/.env. Shell profiles source it, MCP server configs reference ${VAR} from it, and the UI server sources it at startup.
That last part is the catch: the scheduler inherits the env of the server process. Rotate a key in .env, forget the restart, and every scheduled job keeps using the old one until the next reboot.
```bash
# after any .env change
pm2 restart ui-server --update-env
`
A framework-level variant bit me twice: a PHP app with a compiled env cache kept serving a revoked API key long after .env was fixed. A week of 401s in a job that reported success. Now the job's quality gate treats an empty LLM response as a failure.
## IndexNow works, GSC Indexing API doesn't
After every publish or refresh, the job pings search engines.
ChannelResultNotes
---------
IndexNow (Bing, Yandex, Seznam…)200/202, every runkey file at site root, one POST per batch
Google Indexing API403, every runofficially scoped to JobPosting/BroadcastEvent pages
GSC sitemap resubmitworksthe only Google lever that matters for articles
curl -s -X POST https://api.indexnow.org/indexnow \
  -H 'Content-Type: application/json' \
  -d '{"host":"site-a.example","key":"'"$INDEXNOW_KEY"'",
       "urlList":["https://site-a.example/blog/new-slug"]}'

The Indexing API 403 is not a permissions bug to debug. It is Google saying the endpoint is not for blog posts. The job now logs it as expected and moves on instead of retrying.

What I'd change

The numbers in the header hide the honest part: 4 GB of swap is 100% used and the disk is at 89%. Seventeen PM2 apps plus headless Chrome for screenshots plus video renders do not fit comfortably in 15 GiB. Next step is moving rendering and anything Chromium-based to a separate small box, and putting the loopback helper into a shared library instead of the copy-pasted snippet that currently lives in four skills.