Verdicts: 12 | Rejected: 10 | Shipped: 2 | Reviewer: read-only plan mode | Date: 2026-09-30 My content autopilot's weekly site-health pass left me a queue of 12 proposed changes: four title rewrites, three suspected cannibalization pairs, two price-update flags, three new-page drafts. Instead of triaging them myself, I handed the whole queue to a second LLM running headless in plan mode — a mode where it has no write tools and physically cannot edit anything. It rejected 10 of the 12. Every rejection came with a reason I could verify in under a minute from data I already had.
That ratio is the finding. The queue looked like work. It was mostly noise.
The reviewer that can't touch anything
I've already replaced human approval buttons with quality gates on this pipeline, because approve buttons become rubber stamps . But judgment calls — is this title worth changing, are these two pages really cannibalizing — still landed on me, and I was becoming the rubber stamp: twelve proposals, maybe ninety seconds of attention each, a bias toward "sure, ship it" because the proposer sounded confident.
The experiment: delegate the triage. One headless invocation per decision:
# plan mode = browse/read only, no write tools, no side effects
grok --permission-mode plan "$(cat <<'Q'
Queue item: rewrite title + meta of the Canva pricing article.
Context: 9 title/meta edits on the same site are already
running as measured experiments. Current title matches
the query intent.
Required output: APPROVE or REJECT, one reason, max 3 lines.
Q
)" Three properties do the work:
- No write tools. Plan mode grants read and browse only. The reviewer returns a verdict, not a diff. If it "decides" the title needs fixing, nothing happens — the classic failure of an eager agent helping itself to production is structurally removed. Capability restriction beats a prompt that says "don't change anything."
- Different model family. The proposing pipeline runs on two families — an OpenAI-lineage default for generation, a second for the scheduled runs. The reviewer is a third. Self-review agrees with itself: shared training correlates exactly the errors you're trying to catch.
- One decision per question, verdict forced. REJECT or APPROVE, one reason, three lines. Ambiguity is where yes-bias lives; "it depends, but maybe" is not an accepted output.
The 12 verdicts
| Proposal | Evidence in context | Verdict |
| Rewrite title, big-brand article | Query is navigational: 1,236 impr, 0 clicks; tool page owns it; same rewrite already rejected once | REJECT |
| Rewrite 3 more titles | Noise queries, or title already matches intent; 9 measured edits in flight | REJECT ×3 |
| Merge brand article + tool page | Tool page owns the queries; article merge adds nothing | REJECT |
| Merge comparison-article pair | 24 vs 19 impressions — below any sane signal floor | REJECT |
| Link tool page → its pricing article | Pricing article supports the tool page's queries | APPROVE |
| Update Notion AI price | Page serves EUR from European vantage; flag came from a US-locale fetch | NO CHANGE |
| Update Midjourney price | Vendor homepage lists no prices at all | NO CHANGE |
| Draft 3 new gap pages | Price variants of pricing pages that already exist | CANCEL ×3 |
The rules the rejections encoded
The rejections weren't vibes. Each applied a rule the proposing agent knew in theory and ignored in practice.
Navigational queries don't need your article. The shiniest item in the queue targeted a bare brand-name query: 1,236 impressions, zero clicks. The reviewer's read: navigational intent, the tool-directory page owns that query, and rewriting the article title just points a second gun at your own foot. The proposer saw 1,236 impressions and stopped reading. An identical rewrite had been proposed and rejected weeks earlier — I fed that fact into the context, and the proposing pipeline proposed it again anyway. Proposers don't learn; logs do.
Twenty impressions is not a signal. The merge candidate pair sat at 24 vs 19 impressions. Both pages are effectively invisible; "winner takes the URL" on that sample is a coin flip with extra steps, and the merge cost — redirects, link rewrites, reindexing — is real regardless. The floor for acting on a pair should be measured in hundreds of impressions, not tens.
Price diffs need locale pinning. Both "price changed" flags were false. The vendor's page serves EUR to European visitors; the flagging fetcher had seen USD from a US vantage point. This is a whole class of phantom diffs: any unlocalized price scraper generates a steady drip of them, and every one lands in a human queue looking urgent. The fix isn't better review — it's pinning the fetch locale so the diff never fires.
Duplicate intent is cheaper to cancel than to rank. The three gap drafts were all "how much does X cost" variants of pricing pages that already existed. Writing them would have manufactured the exact cannibalization problem the rest of the queue was trying to clean up. Cancelling three drafts cost nothing and closed the loop; the gap finder now skips price variants of existing price pages by rule.
Don't stack edits on running measurements. Nine title/meta changes were mid-measurement when the four new rewrites arrived. Shipping them would have confounded attribution for all thirteen. A reviewer that can see in-flight experiments is worth more than one that reads titles in isolation.
What shipped
Two approvals, both executed the same day. One internal link from a tool page to its matching pricing article. And the single title/meta rewrite the report had pre-approved — registered as a measured change, so its before/after gets tracked like the other nine.
An honest asymmetry: I can't yet prove the two approvals were right — they're being measured. The ten rejections were each checkable on the spot. A reviewer that says no hands you verifiable claims. A reviewer that says yes hands you work.
What I'd change
The reviewer sees SEO context but not action cost — a REJECT on a thirty-second fix and a REJECT on a three-day build read exactly the same, and cheap experiments deserve a lower bar. I'd also rotate the reviewer family between runs: one reviewer has its own blind spots, and overfitting a pipeline to a single model's quirks just relocates the rubber stamp. The part I wouldn't change is the plan-mode constraint. Every other reviewer setup I've tried — same family, full tools, "please only advise" — eventually edits something.