15 runs, 0 confirmed pairs, 24 merges
Period: 2026-07-30 → 2026-09-28 | Runs: 15 | Detector confirmed: 0 | Merged by hand: 24 | Status: detector demoted to a hint list A Czech AI-news site I run publishes and refreshes articles on autopilot. Every refresh run starts with a keyword-cannibalization detector. In two months it confirmed zero cannibalizing pairs. Every single run. Over the same period I merged 24 articles into other articles with a 301, and every one of those merges was real.
The detector wasn't crashing. It was measuring the wrong thing.
How the detector works
Two stages, both reasonable on paper:
- Score every pair of published titles for string similarity (0-100). Keep pairs at 75 or above.
- Pull the top 20 Google Search Console queries for each page in the pair. Confirm cannibalization only if they share 3 or more queries.
Stage 1 produces candidates. Stage 2 is the gate. The gate never opened.
The run log
| Run | Similarity candidates | GSC-confirmed | Merged | How the real pairs were found |
| 07-30 | n/a | 0 | 3 | model-family clustering |
| 08-06 | n/a | 0 | 3 | exact duplicates, gate skipped |
| 08-10 | n/a | 0 | 2 | gate skipped, content review |
| 08-20 | n/a | 0 | 3 | intent + tool-set triage |
| 08-24 | 13 | 0 | 2 | intent + tool-set triage |
| 08-27 | 20 | 0 | 0 | all false positives |
| 08-31 | 10 | 0 | 1 | triage, 0 GSC queries on every candidate |
| 09-03 | n/a | 0 | 2 | grouping recent articles by topic |
| 09-07 | n/a | 0 | 3 | triage, 85% outline overlap on one pair |
| 09-10 | 5 | 0 | 1 | GSC query-split |
| 09-14 | 30 | 0 | 1 | shared queries flagged by a topic-research run |
| 09-17 | 32 | 0 | 0 | 21 of 32 were one price-page cluster |
| 09-21 | n/a | 0 | 1 | triage |
| 09-24 | 32 | 0 | 0 | ~18 price-page false positives |
| 09-28 | 35 | 0 | 2 | cluster triage |
| Total | 0 | 24 | ||
| Three of the 24 were byte-level duplicates (similarity 100, shared slug base) that only surfaced once I ran the detector with the GSC gate disabled. The other 21 were content-level cannibalization that neither stage caught with confidence. | ||||
| ## Failure 1: model-version titles don't look alike | ||||
| The pair that broke my trust was a cluster of four articles about one OpenAI release, GPT-5.4. Same model, same search intent, four framings: computer use, context window, office tasks, the mini variant. Each was thin, 740 to 1,096 words. Each was competing for the same "GPT-5.4" queries. | ||||
| Title similarity for every pair in that cluster: below 75. The headlines share one token that matters and five that don't. String similarity weights all tokens equally, so the one that matters gets outvoted. | ||||
| The fix was consolidation into one winner, refreshed to 1,640 words with GPT-5.6 context, and three 301s. That model line went GPT-5.4 (March) to GPT-5.5 (April) to GPT-5.6 (July). Model-release articles on an AI-news site supersede themselves roughly every three months, which means they cluster by design and the titles drift apart by design. | ||||
| ## Failure 2: templated titles look too alike | ||||
| The opposite problem dominated the candidate list. The site runs a programmatic set of per-tool pricing pages, all titled from one template: "{Tool} pricing and plans 2026". Canva, Zapier, Lovable, Framer, Cursor, ChatGPT, Claude Code. | ||||
| Title similarity: very high. Search intent: completely different, since nobody searching Canva pricing wants Framer pricing. On 09-17, 21 of 32 candidates were this one cluster. On 09-24, about 18 of 32. The detector was spending its entire signal on pages that were supposed to look identical. | ||||
| ## Failure 3: the GSC gate starves on exactly the pages that matter | ||||
| Requiring 3 shared queries from the top 20 sounds conservative. In practice: | ||||
| Fresh duplicates (under a month old) haven't collected enough queries to share 3. They are the ones you want to kill before they split link equity. | ||||
| Stale, low-traffic pages have zero queries. On 08-31, every merge candidate had 0 GSC queries in 90 days. Zero shared queries is not proof of no overlap. It is proof of no data. | ||||
| Pages that do rank usually overlap on 1 or 2 high-volume queries, not 3 long-tail ones. | ||||
| The gate was designed to prevent false positives. It succeeded completely by never producing positives. | ||||
| ## Failure 4: slugs are not titles | ||||
| A smaller trap. On this site slugs don't follow titles. One article's slug ends in a word variant not present in the title; another uses a short workflow slug under a long headline. Anything that tries to join detector output, GSC page URLs and CMS records by deriving one from another breaks silently. Always read the slug field from the CMS. I learned that after submitting two wrong URLs to IndexNow in one run. | ||||
| ## What actually found the 24 | ||||
| ### Cluster by entity, not by string | ||||
| Group published articles by model or tool family: every "GPT-5.x", every "Gemini 3.x", every "Claude" article, every "Cursor vs" article. Then read the thin ones side by side. This found the four-article GPT cluster, a recurring LLM pricing topic that kept regrowing duplicates (a third copy appeared after two were already merged), and an editor-comparison topic that took two merges seven weeks apart to fully consolidate. | ||||
| ### GSC query-split instead of query-overlap | ||||
| Instead of asking "do pages A and B share 3 of their top queries", ask "which queries split impressions across more than one page". Same API, different grouping: | ||||
| ```python | ||||
| from collections import defaultdict | ||||
| def query_splits(gsc, site_url, start, end, min_impr=2): | ||||
| """Queries where Google shows more than one of our pages. | ||||
| Catches pairs that share 1 strong query, which the 3-query gate misses.""" | ||||
| rows = gsc.searchanalytics().query(siteUrl=site_url, body={ | ||||
| "startDate": start, "endDate": end, | ||||
| "dimensions": ["query", "page"], # page per query, not query per page | ||||
| "rowLimit": 25000, | ||||
| }).execute().get("rows", []) | ||||
| by_query = defaultdict(list) | ||||
| for r in rows: | ||||
| query, page = r["keys"] | ||||
| if r["impressions"] >= min_impr: # drop single stray impressions | ||||
| by_query[query].append({ | ||||
| "page": page, | ||||
| "impr": r["impressions"], | ||||
| "pos": round(r["position"], 1), | ||||
| "clicks": r["clicks"], | ||||
| }) | ||||
| splits = {q: sorted(p, key=lambda x: x["pos"]) | ||||
| for q, p in by_query.items() if len(p) > 1} | ||||
| # Google's ranking is the tie-breaker: first page listed is the likely winner | ||||
| return dict(sorted(splits.items(), | ||||
| key=lambda kv: -sum(p["impr"] for p in kv[1]))) | ||||
| ` | ||||
| On 09-10 this surfaced a pair the detector scored as unrelated. Both pages ranked for the same Czech query about creating influencer content: one at position 5.9 with 192 impressions, the other at 11.3 with 22. Google had already picked a winner. On 09-14 it confirmed a pair where the loser had 0 clicks across 3 queries and the winner had 4 clicks across 12. | ||||
| ### Triage criteria that held up across 15 runs | ||||
| Signal | Verdict | |||
| --- | --- | |||
| Same intent + overlapping tool set | merge | |||
| Same template, different product (pricing pages) | false positive | |||
| Different model pairs (Veo vs Sora is not Kling vs Seedance) | false positive | |||
| Specific tutorial vs general guide on a related topic | false positive | |||
| 0 shared GSC queries on pages that have traffic | false positive | |||
| 0 GSC queries on both pages | judge on content alone | |||
| Newer article fully supersedes older one's data | merge, don't port |
The false-positive list matters as much as the merge list. The same rejected pairs came back as candidates run after run. Writing them down stopped the re-triage.
Winner selection: inbound links, then Google, then freshness
The detector picked winners by the direction of the similarity score, which is meaningless. What worked:
- More live internal inbound links wins. Fewer links to rewrite, fewer redirect hops. One beginner-guide merge picked the page with four live inbound links over its twin published the same day.
- If GSC has data, Google's pick wins. Position 5.9 beats 11.3.
- Check which one is actually current. In one editor comparison the "winner" by links was also the one with the newer price hike. The loser still quoted the old $15 plan.
Fix the winner during the merge
A merge is a refresh. Across the 24:
- One winner still had a model table from the GPT-4o era. Replaced with the current lineup during the merge.
- One loser claimed a product was in limited preview. It had gone GA in June 2026. The merged winner got the correct status.
- In one merge, the loser had the better deployment and ROI sections. Those were grafted in instead of trusting the winner body.
- On 09-28 the loser's data was entirely superseded, so it got a 301 with nothing ported.
- One winner's intro linked to the loser's slug. After the 301 that would have been a redirect loop. Grep the winner body for the loser slug before publishing.
Every 301 got verified with a real request, not assumed from the CMS response.
What I'd change
Title similarity should never have been the primary signal. It is now a hint list for spotting exact duplicates at 100 and nothing else. The actual detector should be three passes: entity clustering (model/tool family extracted from titles and bodies), GSC query-split with a 1-query threshold weighted by impressions, and a persistent false-positive list so the same price pages stop resurfacing. I'd also exclude the programmatic pricing template from pairwise scoring entirely. It generated more noise than the other 300+ articles combined.