← Back to lab

83% of My SEO To-Do List Was Noise. A Read-Only Model Proved It.

My content autopilot queued 12 SEO changes. A second LLM, run headless in plan mode with no write tools, rejected 10 of them — each for a checkable reason.

83% of My SEO To-Do List Was Noise. A Read-Only Model Proved It.
Verdicts: 12  |  Rejected: 10  |  Shipped: 2  |  Reviewer: read-only plan mode  |  Date: 2026-09-30

My content autopilot's weekly site-health pass left me a queue of 12 proposed changes: four title rewrites, three suspected cannibalization pairs, two price-update flags, three new-page drafts. Instead of triaging them myself, I handed the whole queue to a second LLM running headless in plan mode — a mode where it has no write tools and physically cannot edit anything. It rejected 10 of the 12. Every rejection came with a reason I could verify in under a minute from data I already had.

That ratio is the finding. The queue looked like work. It was mostly noise.

The reviewer that can't touch anything

I've already replaced human approval buttons with quality gates on this pipeline, because approve buttons become rubber stamps . But judgment calls — is this title worth changing, are these two pages really cannibalizing — still landed on me, and I was becoming the rubber stamp: twelve proposals, maybe ninety seconds of attention each, a bias toward "sure, ship it" because the proposer sounded confident.

The experiment: delegate the triage. One headless invocation per decision:

# plan mode = browse/read only, no write tools, no side effects
grok --permission-mode plan "$(cat <<'Q'
Queue item: rewrite title + meta of the Canva pricing article.
Context: 9 title/meta edits on the same site are already
running as measured experiments. Current title matches
the query intent.
Required output: APPROVE or REJECT, one reason, max 3 lines.
Q
)"

Three properties do the work:

  • No write tools. Plan mode grants read and browse only. The reviewer returns a verdict, not a diff. If it "decides" the title needs fixing, nothing happens — the classic failure of an eager agent helping itself to production is structurally removed. Capability restriction beats a prompt that says "don't change anything."
  • Different model family. The proposing pipeline runs on two families — an OpenAI-lineage default for generation, a second for the scheduled runs. The reviewer is a third. Self-review agrees with itself: shared training correlates exactly the errors you're trying to catch.
  • One decision per question, verdict forced. REJECT or APPROVE, one reason, three lines. Ambiguity is where yes-bias lives; "it depends, but maybe" is not an accepted output.

The 12 verdicts

ProposalEvidence in contextVerdict
Rewrite title, big-brand articleQuery is navigational: 1,236 impr, 0 clicks; tool page owns it; same rewrite already rejected onceREJECT
Rewrite 3 more titlesNoise queries, or title already matches intent; 9 measured edits in flightREJECT ×3
Merge brand article + tool pageTool page owns the queries; article merge adds nothingREJECT
Merge comparison-article pair24 vs 19 impressions — below any sane signal floorREJECT
Link tool page → its pricing articlePricing article supports the tool page's queriesAPPROVE
Update Notion AI pricePage serves EUR from European vantage; flag came from a US-locale fetchNO CHANGE
Update Midjourney priceVendor homepage lists no prices at allNO CHANGE
Draft 3 new gap pagesPrice variants of pricing pages that already existCANCEL ×3

The rules the rejections encoded

The rejections weren't vibes. Each applied a rule the proposing agent knew in theory and ignored in practice.

Navigational queries don't need your article. The shiniest item in the queue targeted a bare brand-name query: 1,236 impressions, zero clicks. The reviewer's read: navigational intent, the tool-directory page owns that query, and rewriting the article title just points a second gun at your own foot. The proposer saw 1,236 impressions and stopped reading. An identical rewrite had been proposed and rejected weeks earlier — I fed that fact into the context, and the proposing pipeline proposed it again anyway. Proposers don't learn; logs do.

Twenty impressions is not a signal. The merge candidate pair sat at 24 vs 19 impressions. Both pages are effectively invisible; "winner takes the URL" on that sample is a coin flip with extra steps, and the merge cost — redirects, link rewrites, reindexing — is real regardless. The floor for acting on a pair should be measured in hundreds of impressions, not tens.

Price diffs need locale pinning. Both "price changed" flags were false. The vendor's page serves EUR to European visitors; the flagging fetcher had seen USD from a US vantage point. This is a whole class of phantom diffs: any unlocalized price scraper generates a steady drip of them, and every one lands in a human queue looking urgent. The fix isn't better review — it's pinning the fetch locale so the diff never fires.

Duplicate intent is cheaper to cancel than to rank. The three gap drafts were all "how much does X cost" variants of pricing pages that already existed. Writing them would have manufactured the exact cannibalization problem the rest of the queue was trying to clean up. Cancelling three drafts cost nothing and closed the loop; the gap finder now skips price variants of existing price pages by rule.

Don't stack edits on running measurements. Nine title/meta changes were mid-measurement when the four new rewrites arrived. Shipping them would have confounded attribution for all thirteen. A reviewer that can see in-flight experiments is worth more than one that reads titles in isolation.

What shipped

Two approvals, both executed the same day. One internal link from a tool page to its matching pricing article. And the single title/meta rewrite the report had pre-approved — registered as a measured change, so its before/after gets tracked like the other nine.

An honest asymmetry: I can't yet prove the two approvals were right — they're being measured. The ten rejections were each checkable on the spot. A reviewer that says no hands you verifiable claims. A reviewer that says yes hands you work.

What I'd change

The reviewer sees SEO context but not action cost — a REJECT on a thirty-second fix and a REJECT on a three-day build read exactly the same, and cheap experiments deserve a lower bar. I'd also rotate the reviewer family between runs: one reviewer has its own blind spots, and overfitting a pipeline to a single model's quirks just relocates the rubber stamp. The part I wouldn't change is the plan-mode constraint. Every other reviewer setup I've tried — same family, full tools, "please only advise" — eventually edits something.