← Back to lab

My Cannibalization Detector Found 0 Duplicates. I Merged 24.

A title-similarity cannibalization detector flagged 0 pairs in 15 runs on my automated content site. Manual triage merged 24 pages. Here is why it failed.

My Cannibalization Detector Found 0 Duplicates. I Merged 24.

15 runs, 0 confirmed pairs, 24 merges

Period: 2026-07-30 → 2026-09-28  |  Runs: 15  |  Detector confirmed: 0  |  Merged by hand: 24  |  Status: detector demoted to a hint list

A Czech AI-news site I run publishes and refreshes articles on autopilot. Every refresh run starts with a keyword-cannibalization detector. In two months it confirmed zero cannibalizing pairs. Every single run. Over the same period I merged 24 articles into other articles with a 301, and every one of those merges was real.

The detector wasn't crashing. It was measuring the wrong thing.

How the detector works

Two stages, both reasonable on paper:

  1. Score every pair of published titles for string similarity (0-100). Keep pairs at 75 or above.
  2. Pull the top 20 Google Search Console queries for each page in the pair. Confirm cannibalization only if they share 3 or more queries.

Stage 1 produces candidates. Stage 2 is the gate. The gate never opened.

The run log

RunSimilarity candidatesGSC-confirmedMergedHow the real pairs were found
07-30n/a03model-family clustering
08-06n/a03exact duplicates, gate skipped
08-10n/a02gate skipped, content review
08-20n/a03intent + tool-set triage
08-241302intent + tool-set triage
08-272000all false positives
08-311001triage, 0 GSC queries on every candidate
09-03n/a02grouping recent articles by topic
09-07n/a03triage, 85% outline overlap on one pair
09-10501GSC query-split
09-143001shared queries flagged by a topic-research run
09-17320021 of 32 were one price-page cluster
09-21n/a01triage
09-243200~18 price-page false positives
09-283502cluster triage
Total024
Three of the 24 were byte-level duplicates (similarity 100, shared slug base) that only surfaced once I ran the detector with the GSC gate disabled. The other 21 were content-level cannibalization that neither stage caught with confidence.
## Failure 1: model-version titles don't look alike
The pair that broke my trust was a cluster of four articles about one OpenAI release, GPT-5.4. Same model, same search intent, four framings: computer use, context window, office tasks, the mini variant. Each was thin, 740 to 1,096 words. Each was competing for the same "GPT-5.4" queries.
Title similarity for every pair in that cluster: below 75. The headlines share one token that matters and five that don't. String similarity weights all tokens equally, so the one that matters gets outvoted.
The fix was consolidation into one winner, refreshed to 1,640 words with GPT-5.6 context, and three 301s. That model line went GPT-5.4 (March) to GPT-5.5 (April) to GPT-5.6 (July). Model-release articles on an AI-news site supersede themselves roughly every three months, which means they cluster by design and the titles drift apart by design.
## Failure 2: templated titles look too alike
The opposite problem dominated the candidate list. The site runs a programmatic set of per-tool pricing pages, all titled from one template: "{Tool} pricing and plans 2026". Canva, Zapier, Lovable, Framer, Cursor, ChatGPT, Claude Code.
Title similarity: very high. Search intent: completely different, since nobody searching Canva pricing wants Framer pricing. On 09-17, 21 of 32 candidates were this one cluster. On 09-24, about 18 of 32. The detector was spending its entire signal on pages that were supposed to look identical.
## Failure 3: the GSC gate starves on exactly the pages that matter
Requiring 3 shared queries from the top 20 sounds conservative. In practice:
Fresh duplicates (under a month old) haven't collected enough queries to share 3. They are the ones you want to kill before they split link equity.
Stale, low-traffic pages have zero queries. On 08-31, every merge candidate had 0 GSC queries in 90 days. Zero shared queries is not proof of no overlap. It is proof of no data.
Pages that do rank usually overlap on 1 or 2 high-volume queries, not 3 long-tail ones.
The gate was designed to prevent false positives. It succeeded completely by never producing positives.
## Failure 4: slugs are not titles
A smaller trap. On this site slugs don't follow titles. One article's slug ends in a word variant not present in the title; another uses a short workflow slug under a long headline. Anything that tries to join detector output, GSC page URLs and CMS records by deriving one from another breaks silently. Always read the slug field from the CMS. I learned that after submitting two wrong URLs to IndexNow in one run.
## What actually found the 24
### Cluster by entity, not by string
Group published articles by model or tool family: every "GPT-5.x", every "Gemini 3.x", every "Claude" article, every "Cursor vs" article. Then read the thin ones side by side. This found the four-article GPT cluster, a recurring LLM pricing topic that kept regrowing duplicates (a third copy appeared after two were already merged), and an editor-comparison topic that took two merges seven weeks apart to fully consolidate.
### GSC query-split instead of query-overlap
Instead of asking "do pages A and B share 3 of their top queries", ask "which queries split impressions across more than one page". Same API, different grouping:
```python
from collections import defaultdict
def query_splits(gsc, site_url, start, end, min_impr=2):
"""Queries where Google shows more than one of our pages.
Catches pairs that share 1 strong query, which the 3-query gate misses."""
rows = gsc.searchanalytics().query(siteUrl=site_url, body={
"startDate": start, "endDate": end,
"dimensions": ["query", "page"], # page per query, not query per page
"rowLimit": 25000,
}).execute().get("rows", [])
by_query = defaultdict(list)
for r in rows:
query, page = r["keys"]
if r["impressions"] >= min_impr: # drop single stray impressions
by_query[query].append({
"page": page,
"impr": r["impressions"],
"pos": round(r["position"], 1),
"clicks": r["clicks"],
})
splits = {q: sorted(p, key=lambda x: x["pos"])
for q, p in by_query.items() if len(p) > 1}
# Google's ranking is the tie-breaker: first page listed is the likely winner
return dict(sorted(splits.items(),
key=lambda kv: -sum(p["impr"] for p in kv[1])))
`
On 09-10 this surfaced a pair the detector scored as unrelated. Both pages ranked for the same Czech query about creating influencer content: one at position 5.9 with 192 impressions, the other at 11.3 with 22. Google had already picked a winner. On 09-14 it confirmed a pair where the loser had 0 clicks across 3 queries and the winner had 4 clicks across 12.
### Triage criteria that held up across 15 runs
SignalVerdict
------
Same intent + overlapping tool setmerge
Same template, different product (pricing pages)false positive
Different model pairs (Veo vs Sora is not Kling vs Seedance)false positive
Specific tutorial vs general guide on a related topicfalse positive
0 shared GSC queries on pages that have trafficfalse positive
0 GSC queries on both pagesjudge on content alone
Newer article fully supersedes older one's datamerge, don't port

The false-positive list matters as much as the merge list. The same rejected pairs came back as candidates run after run. Writing them down stopped the re-triage.

Winner selection: inbound links, then Google, then freshness

The detector picked winners by the direction of the similarity score, which is meaningless. What worked:

  1. More live internal inbound links wins. Fewer links to rewrite, fewer redirect hops. One beginner-guide merge picked the page with four live inbound links over its twin published the same day.
  2. If GSC has data, Google's pick wins. Position 5.9 beats 11.3.
  3. Check which one is actually current. In one editor comparison the "winner" by links was also the one with the newer price hike. The loser still quoted the old $15 plan.

Fix the winner during the merge

A merge is a refresh. Across the 24:

  • One winner still had a model table from the GPT-4o era. Replaced with the current lineup during the merge.
  • One loser claimed a product was in limited preview. It had gone GA in June 2026. The merged winner got the correct status.
  • In one merge, the loser had the better deployment and ROI sections. Those were grafted in instead of trusting the winner body.
  • On 09-28 the loser's data was entirely superseded, so it got a 301 with nothing ported.
  • One winner's intro linked to the loser's slug. After the 301 that would have been a redirect loop. Grep the winner body for the loser slug before publishing.

Every 301 got verified with a real request, not assumed from the CMS response.

What I'd change

Title similarity should never have been the primary signal. It is now a hint list for spotting exact duplicates at 100 and nothing else. The actual detector should be three passes: entity clustering (model/tool family extracted from titles and bodies), GSC query-split with a 1-query threshold weighted by impressions, and a persistent false-positive list so the same price pages stop resurfacing. I'd also exclude the programmatic pricing template from pairwise scoring entirely. It generated more noise than the other 300+ articles combined.