Back to blog
NO.
021
DATE
READ
~25 min
KIND
Guide
STATUS
Reviewed

TAGS: SEO

Multi-Site SEO Governance: Clusters to 301s

With a dozen English sites, the conflict is not duplicate keywords — it is that no domain, page or person owns the demand. A framework from GSC to 301.

This is not a "site network SEO" guide. In Chinese, that phrase usually means PBNs — fake sites built to link at each other — bulk site generation, and search manipulation. What this post addresses is an entirely legitimate problem:

A company runs a dozen real business websites — different brands, markets, product lines. How should search demand be divided between them, how should duplicate pages be handled, and how do you measure the result in inquiries rather than rankings?

Three kinds of content run through this post, each tagged:

  • [OFFICIAL]: rules written down in Google's own documentation. Every one carries a source link you can check.
  • [FRAMEWORK]: methods I derived from those rules and from delivery experience. Usable as-is, but not the only right answer — if your data says otherwise, follow your data.
  • [EXAMPLE]: a fictional B2B plastic pallet company, used to keep the process concrete. Every number in it is invented and represents no real client.

Conclusions first

  • The root cause of conflict is not "several pages use the same keyword." It is that a single search demand has no assigned page to serve it. The result is several pages competing for the same buyers while nobody can say which page produced the inquiry.
  • [OFFICIAL] Google does not penalize you for duplicate content unless you are deliberately manipulating search. What it does is group duplicate URLs, pick one of them to show, and consolidate link signals onto that one. The catch: Google picks, not you.
  • [OFFICIAL] There is one hard line for multiple domains. Google lists "multiple websites with slight variations to the URL and home page" built to maximize coverage of a query as doorway abuse. A legitimate multi-site setup requires independent value per site, not a new domain over the same content.
  • The fix is not to merge everything immediately. It is: assign a primary page to high-value demand first, then differentiate and freeze, and only then consider a cross-domain 301.
  • "Position 4–30" is never a filter on its own. With 1 impression and 0 clicks you cannot even claim a low CTR — see the statistics in "Finding the top 50 clusters in GSC."
  • A central keyword table works, but only when it is far smaller than most teams imagine: 20–100 high-value clusters, plus a publishing gate, plus one person with decision rights. Missing any of the three and it becomes an unmaintained document within three months.

Every hard claim in this post traces back here:

ClaimKey factOfficial source
Near-duplicate sites are doorway abuse"multiple websites with slight variations to the URL and home page"Google Spam Policies
How long to keep a 301"as long as possible, generally at least 1 year"Site move with URL changes
Canonical is a signal, not an orderredirect (strong) > rel=canonical (strong) > sitemap (weak)Consolidate duplicate URLs
302 does not pass canonicalizationA temporary redirect is not a signal that the target should be canonicalRedirects and Google Search
GSC data ceilingsAPI: 50K rows/day/search type/property; rowLimit ≤ 25000 per requestExport data using the API
Bulk export does not backfillFirst export within 48 hours of setup; earlier data excludedBulk data export
Average position in BigQuerysum_position is zero-based: SUM(sum_position)/SUM(impressions) + 1Export table fields
Limits of the site-move toolDomain-level only, both properties verified, 180 days of forwardingChange of address tool
What GA4 counts as a landing pagePath of the session's first pageview, session-scopedLanding page report

Why more sites is not more search coverage

Running several sites is normal. At least four reasons are legitimate: separate brands, separate languages, separate legal entities, and separate product lines serving non-overlapping buyers. The problem is the fifth reason — "one more site means one more stream of traffic."

Three conflicts that are not the same thing

Teams routinely conflate them, and the treatments differ completely:

TypeDefinitionWhere it happensPolicy violationDirection
On-site cannibalizationSeveral pages on one site compete for the same demandWithin one domainNoMerge or differentiate
Cross-domain competitionSeveral domains of the same company compete for one demandAcross owned domainsNoDivide, freeze, merge if necessary
Doorway abuseNear-duplicate sites or pages built to capture rankingsAcross domains or pagesYesStop and remediate

The first two are efficiency problems, not violations. [OFFICIAL] Google's position is that duplicate content is not itself penalized unless it is deceptive. What Google does is group the duplicate URLs, choose a representative, and consolidate link signals onto it.

The third is a violation. [OFFICIAL] The spam policies name three doorway patterns: building a set of sites that differ only slightly in URL and homepage to capture a query; building numerous regional domains that funnel everyone to one page; and generating pages en masse purely as entry points into the useful part of a site. The same policy also treats spreading mass-produced content across multiple sites to disguise its scale as scaled content abuse.

Three questions to test your own line

[FRAMEWORK] Ask them of every site from the second onward:

  1. Independent value: if this site went dark, would customers lose something they cannot get elsewhere — a product line, a service scope, a legal entity, local support?
  2. Independent operation: does it have its own product data, quoting process, and contacts, or is it the same content on a different template and domain?
  3. Independent demand: can you state the search demand it serves in one sentence, and is that sentence different from every other site's?

A site that fails all three does not need SEO. It needs a decision: shut down, absorb, or reposition.

Action for this section: list every owned domain and answer these three questions in a table. That table is the Domain Charter referenced in "Can a central keyword table survive contact with reality."


From query to inquiry: the whole decision chain

Most multi-site plans fail because they only do the first half — split the keywords across sites, then stop. The full chain is a loop:

The eight-step query-to-inquiry loop: real query, search intent, keyword cluster, primary domain, primary page, supporting content, landing-page inquiry, sales feedback — and sales feedback flowing back to redraw cluster edges, re-rank priorities and reassign pages

Everything depends on that last return line: inquiries and sales feedback change the cluster's boundaries, its page assignment, and its priority. Without it, the first seven steps are a one-off keyword split that drifts out of reality within a quarter.

[EXAMPLE] The company used throughout

Fictional company P manufactures plastic pallets. It runs 12 English sites, 4 of them genuinely active.

① Real query — what people actually type:

food grade plastic pallets
hygienic plastic pallet supplier
sanitary pallets for food industry

Queries are raw data. Do not build one page per query — that discipline is the prerequisite for everything below.

② Search intent — three different wordings, one job: find a hygienic plastic pallet supplier for food, pharma, or cleanroom use. Procurement intent, not research.

③ Keyword cluster — group the queries one page can satisfy, and you get Cluster: Hygienic / Food-grade Plastic Pallet. The unit of a cluster is a demand, not a fixed phrase.

④ Primary domain — only one of the 12 sites should carry it: site B, which serves industrial pallet buyers. The decision weighs target countries, positioning, existing rankings, backlink quality, product documentation, and inquiries already won. It does not go to whoever ranked for the term first.

⑤ Primary page — one procurement page per cluster, /hygienic-plastic-pallet/, carrying the product description, specification table, applications, quality evidence (material certificates and certifications), and the quote request.

⑥ Supporting pages — not copies of the primary page, but answers to narrower questions: food industry case studies, cleaning and sanitation guidance, HDPE vs. PP comparison, cold storage and pharma warehousing, sizing and load selection. They link into the primary page — and must never title themselves Hygienic Plastic Pallet Supplier. That is the single most common origin of cannibalization, usually created by a content team "adding a bit of SEO."

⑦⑧ Connecting to inquiries — the real data chain is not "one keyword equals one customer":

GSC:        query → landing page
GA4:        landing page → inquiry key event
CRM/sales:  inquiry → qualified / disqualified / quoted / won

[OFFICIAL] In GA4 the landing page is the first page opened in a session. Key events appear in the same report — but a key event requires you to collect the event and then mark it as key; it does not appear on its own.

One limit you must state out loud: privacy protection and data truncation mean you generally cannot map each query to a specific customer. You approximate at the level of landing page, product, country, and period. Any vendor claiming precise keyword-level attribution for organic inquiries should be asked how they get around GSC's anonymized queries.

Action for this section: in your analysis sheet, keep "keyword" and "landing page" as two separate columns, and attach conversion data only to the landing page column. This one rule blocks about 80% of downstream attribution errors.


What a keyword cluster is, and is not

The same words are not always the same intent: plastic pallet might be a buyer, a price check, or a student. The tell is not in the words but in what kinds of pages rank on the SERP.

Different words are often the same intent: hygienic plastic pallet and sanitary pallets for food industry share no vocabulary and are satisfied by the same page.

So the grouping test is "can one page satisfy all of these," not "do these look alike."

Five criteria

[FRAMEWORK] To decide whether two queries belong to one cluster, walk all five:

DimensionThe question[EXAMPLE]
Product entityDo they point at the same sellable item?Hygienic plastic pallet = one entity
Nature of modifiersDoes the modifier change an attribute or the product?food grade / hygienic / sanitary are all one attribute: cleanliness grade
Market and languageSame country and language?English / US, UK
Intent typeBuy, compare, learn, or navigate?All procurement
SERP overlapHow many of the top 10 URLs are shared?See below

How to measure SERP overlap: take 3–5 representative queries from the cluster, search each once in the target country in an incognito window, and record the top 10 organic URLs. [FRAMEWORK] Three or more shared URLs → treat as one cluster; one or fewer → separate clusters; exactly two → manual review pool. That threshold is my working number, not a Google standard; adjust it for your vertical, but fix it before you start. An unfixed threshold means grouping is decided by whoever argues hardest.

One positive and one negative example

[EXAMPLE] Positive — one cluster:

food grade plastic pallets
hygienic plastic pallet supplier
sanitary pallets for food industry
plastic pallets for food processing

The top 10 results overlap heavily and are all supplier product and category pages. One /hygienic-plastic-pallet/ serves all four.

[EXAMPLE] Negative — looks like one cluster, must be split:

plastic pallet price per unit          → procurement, price-comparison stage
plastic pallet vs wooden pallet        → comparison, selection stage
how to clean plastic pallets           → usage, existing customer
plastic pallet recycling               → disposal, an entirely different journey

All four contain plastic pallet and sit at different stages, with different result types (price pages, comparison articles, how-to guides, recycling services). Force them onto one page and none of them rank.

When one cluster needs both an informational and a procurement page

[FRAMEWORK] When the top 10 for a cluster consistently contains two page types (say 6 product pages and 4 tutorials), Google is telling you the query carries two demands. The response: keep exactly one primary page for procurement intent, build a separate informational page for the learning intent, link them, and point the informational page's call to action at the primary page. The two titles must be distinguishable at a glance by someone who does not know the business.

Action for this section: write one sentence per candidate cluster stating what question its page answers. If you cannot write that sentence, it is not yet a cluster.


Finding the top 50 clusters worth governing in GSC

Pick the right extraction method

A dozen sites cannot be copied out by hand. The three methods have different ceilings:

MethodPer-request ceilingHistoryScale it suitsOfficial basis
GSC UI export1,000 rowsYesSpot checks on one site—
Search Analytics APIrowLimit ≤ 25000/request, 50K rows/day/type/propertyYes3–20 sitesAPI reference, export limits
BigQuery bulk exportNo practical row limitNot backfilledLong-term build-outBulk data export

Here is the trap. [OFFICIAL] The first batch of a BigQuery bulk export starts accumulating within 48 hours of a successful setup. Data from before setup is not available, and Google says explicitly to use the API or the UI report for history. So the correct order is:

  1. Configure bulk export for every site today — it only accrues forward, so every day of delay is a day of data lost.
  2. In parallel, pull six months of history through the Search Analytics API (startRow paging, rowLimit 25000).
  3. Store the two datasets separately and do not try to reconcile them — their aggregation differs, as below.

One aggregation difference you must know

[OFFICIAL] GSC aggregates differently per dimension: query, country, device, and date aggregate by property; page and search appearance aggregate by page.

That is why the chart totals and the table totals in GSC so often disagree.

What it means for you: you cannot sum per-page query impressions and call it the cluster's impressions. One search that surfaced several of your URLs is counted differently under the two methods. [FRAMEWORK] My rule: read cluster-level impressions only from query-dimension totals, and use page-dimension data only to answer "which page is this query landing on." Report the two numbers in separate columns and never add or subtract across them.

The correct way to compute average position in BigQuery

[OFFICIAL] The sum_position field in the export is zero-based (position 1 is stored as 0), so average position needs a +1 at the end:

-- searchdata_url_impression: cluster performance by page
SELECT
  url,
  SUM(impressions)                                     AS impressions,
  SUM(clicks)                                          AS clicks,
  SAFE_DIVIDE(SUM(clicks), SUM(impressions))           AS ctr,
  SAFE_DIVIDE(SUM(sum_position), SUM(impressions)) + 1 AS avg_position
FROM `project.searchconsole.searchdata_url_impression`
WHERE data_date BETWEEN '2026-01-01' AND '2026-06-30'
  AND country = 'usa'
  AND is_anonymized_query = FALSE
  AND REGEXP_CONTAINS(query, r'(hygienic|sanitary|food.grade).*pallet')
GROUP BY url
ORDER BY impressions DESC

Writing AVG(position) is wrong: it is unweighted and will drift from the real average.

The seven steps

[FRAMEWORK]

  1. Export six months of query, page, country, clicks, impressions, CTR, and position for every site.
  2. Strip branded queries. [OFFICIAL] Since 2025-03-11 the performance report offers a branded/non-branded filter, but it does not apply to sub-properties (like example.com/blog/) or to sites with too few impressions — if you do not see it, maintain your own brand regex (including misspellings and domain variants). Use the same regex when pulling through the API.
  3. Normalize — for candidate grouping only, never overwriting the raw query column: lowercase, trim, collapse whitespace; singularize (pallets → pallet); drop stop words (for, of, the, in); sort the words into a key so that plastic pallets food grade and food grade plastic pallet both become food grade pallet plastic; and maintain a synonym map by hand (hygienic ↔ sanitary ↔ food grade). That last step cannot be automated, because whether two words are synonyms depends on the industry.
  4. Group by product + intent, then verify with the SERP overlap test above. Normalization produces candidates; a cluster is confirmed by a human.
  5. Aggregate impressions, clicks, and average position across the cluster (minding the aggregation difference above).
  6. Add two non-GSC columns: commercial value of the product (filled in by sales or product management, not by SEO) and qualified inquiries from that cluster's landing page over the last six months.
  7. Score, rank, and take the top 20–50 into governance.

The high-value cluster scorecard

[FRAMEWORK] Out of 100. This is a sorting tool, not a predictor — its job is to make "which one first" defensible, not to promise that a high score performs.

DimensionWeightScoring
Demand size D25Cluster monthly impressions ≥5000 → 25; 1000–4999 → 18; 200–999 → 12; 50–199 → 6; <50 → 2
Commercial value V30Sales scores 1–5 on margin × repeat purchase × cycle length, then ×6
Competitive feasibility F20Own best position 1–3 → 8 (already held, little headroom); 4–10 → 20; 11–30 → 14; 31+ or absent → 6
Asset readiness R15Page exists with full specs and evidence → 15; page but no evidence → 9; neither → 3
Conflict cost C10Own pages competing for the cluster ≥4 → 10; 2–3 → 6; 1 → 0

One override: if a cluster produced even one sales-confirmed qualified inquiry in the last six months, it goes into the manual review pool regardless of score. Evidence of a real deal outranks a big number.

[EXAMPLE] Hygienic pallet cluster: D=18 (2,400 monthly impressions), V=24 (sales scored 4), F=20 (best position 7), R=9 (page exists, certificates missing), C=6 (three pages competing) → 77, into the top 10.

What to do with "one impression"

This is the most abused number in practice. Start with the intuition: the fewer observations you have, the less "it did not happen" is allowed to be a conclusion.

Statistics has a shortcut called the Rule of Three: if you observed n times and the event never happened, the most you can say is that its probability is roughly at most 3 ÷ n.

Concretely: flip a coin 100 times with no heads, and 3 ÷ 100 = 3% — you can say with reasonable confidence that heads is under 3%. Flip it once, and 3 ÷ 1 = 300%. Since 300% exceeds 100%, you have concluded nothing at all.

Applied to impressions and clicks:

ImpressionsWith 0 clicks, the true CTR could be as high asWhat it tells you
1~300% (higher than 100%, i.e. meaningless)Nothing
30~10%Almost nothing
100~3%Barely "CTR is under 3%"
300~1%You may say "CTR really is low"

In short: to claim "this cluster's CTR is abnormally low," you need hundreds of impressions. And "low" must be measured against your own site's historical CTR at comparable positions, never an industry average. By the same logic, an "average position" attached to a single impression is one isolated number that cannot support any trend.

[OFFICIAL] There is a further complication: queries with very low volume are anonymized and disappear from the report entirely. So "only 1 impression" does not even prove there was only 1 impression — others may be hidden.

[FRAMEWORK] Six situations, six treatments:

SituationTreatment
1 impression, 0 clicks, 0 inquiriesObservation pool. Do not change a page over it
1 impression a month for several monthsStill too sparse. Keep observing; usable as a topic idea
20 near-synonym queries at 1 impression eachWorth reading together (20 impressions is still thin, but it is at least one unit of observation)
1 impression → 1 click → 1 qualified inquirySparse data, high commercial value: verify the inquiry's real source by hand
Sales says the product is valuableTest content without waiting for impressions to accumulate
Position 4–30 but under 50 impressionsNot enough to decide on. Back to the observation pool

So "position 4–30" can never be a filter by itself. It has to be read alongside impressions, recurrence, commercial value, and inquiries.

Action for this section: configure bulk export for every site today (it does not backfill — a day late is a day lost), pull six months of history through the API in parallel, then produce the top 20 clusters with the scorecard.


Two pages: differentiate or merge

The decision tree

[FRAMEWORK] Ask in order and stop at the first match. Do not skip ahead:

Do both pages hit the same cluster?
├─ No → not a conflict. Register as separate clusters. Done.
└─ Yes
   ├─ Q1 Same search intent?
   │      No → [KEEP BOTH] state each intent, rewrite titles, cross-link
   ├─ Q2 Same target country or language?
   │      No → [KEEP BOTH] add hreflang. Never 301
   ├─ Q3 Different product entity, specification, or application?
   │      Yes → [KEEP BOTH] give each its own cluster
   ├─ Q4 Did both produce qualified inquiries in the last 6 months?
   │      Yes → [DIFFERENTIATE] do not merge a page with real inquiries for tidiness
   ├─ Q5 Does the weaker page hold its own backlinks, brand equity, or partner citations?
   │      Yes → prefer [DIFFERENTIATE]; if you must merge, the 301 is what preserves those links
   └─ All no → [MERGE + 301]

Differentiation is not a title rewrite. It is four things at once: titles and H1s (so a reader can tell the two pages apart instantly); the first screen of body copy (the first 200 words say who this page is for); internal link direction (weak page points at strong page, never mutually); and the call to action (one asks for a quote, the other offers a selection guide download).

[EXAMPLE] The seven steps of a merge

Company P's site B has two pages fighting over the hygienic pallet cluster: /hygienic-plastic-pallet/ (1,800 monthly impressions, 5 inquiries) and /food-grade-pallet-supplier/ (260 impressions, 0 inquiries, no backlinks). The decision tree lands on merge + 301.

  1. Back up both pages: 12 months of GSC performance each, GA4 landing page data, a snapshot of the current pages, and the backlink list. This is your rollback basis and the baseline you will judge the merge against.
  2. Move the content: re-edit the weak page's unique specifications, FAQs, images, and cases into the strong page. Not copy-paste — integrate into the strong page's structure.
  3. Rewrite internal links: change every site-wide link pointing at the weak URL to the strong URL. Do not leave it to the 301 — every extra hop is another chance to break.
  4. Ship the 301: weak URL → strong URL, server-side 301 or 308. [OFFICIAL] Google states server-side redirects are "the most reliably detected," and a temporary redirect (302) does not signal that the target should become canonical.
  5. Update sitemap, canonical, and structured data: drop the weak URL from the sitemap; keep the strong page's self-referencing canonical; merge Product/FAQ structured data into the strong page.
  6. Eliminate redirect chains: verify there is no A→B→C hop; make everything A→C.
  7. Monitor: [OFFICIAL] use Search Console's sitemap, indexing, and performance reports together to watch crawling and traffic on both old and new URLs. Give it 4–8 weeks. Do not conclude anything on day five.

Action for this section: turn step 1 into a template file (the companion toolkit has the full checklist). A merge without baseline data can neither be judged nor rolled back.


Why a 301 stays up for at least a year

[OFFICIAL] Google's site move documentation says it plainly:

Keep the redirects for as long as possible, generally at least 1 year.

One common misreading needs correcting: "at least a year" is a floor, not permission to delete on day 366.

Four reasons:

  1. [OFFICIAL] A 301 is a signal that the URL moved, not an instruction that executes immediately. Google has to recrawl both URLs before the transfer completes.
  2. Crawl frequency varies wildly per page. A low-traffic page may be visited once every few months.
  3. Old links live far longer than a year — bookmarks, emails from years ago, PDF quotations saved by customers, third-party B2B directories, trade association listings.
  4. Pull the 301 early and all of those become 404s at once. What you lose is not only a search signal; it is buyers who were trying to reach you.

The 180 days that get confused with it

[OFFICIAL] The Change of address tool forwards signals from the old site to the new one — for 180 days.

Those 180 days are not the lifespan of your 301. They are the window in which Google actively forwards the signal. After it closes, Google no longer treats the two sites as related, and your redirects still need to stay up: URLs remain uncrawled, and humans keep using old links.

Confusing those two numbers is one of the most expensive migration errors I have seen.

Three situations, three answers

[FRAMEWORK]

SituationRecommendationWhy
Whole-site moveAt least a year — the official floorEvery signal has to be redistributed
Single page mergeKeep indefinitely if maintenance is cheapOne rewrite rule costs nothing; the payoff is old links that never break
Content is gone with no suitable replacementReturn 404 or 410. Do not redirect to the homepageRedirecting unrelated URLs to the homepage is noise for users and for Google

The third deserves emphasis. Do not bulk-redirect retired pages to the homepage to "avoid 404s." Users who arrive and find something else simply leave, and Google treats that pattern as a soft 404. When no equivalent content exists, an honest 410 ("permanently gone") works better than a lazy homepage redirect.

[REAL] A small example that only proves single-site process

When I moved this blog from yesyes.qzz.io to eigentime.org, I decided to keep the old domain's 301s for at least 12 months as a rollback guarantee; the full retrospective is here. I apply the same discipline within the site: path change → add a 301; retracted post → return 410 plus X-Robots-Tag: noindex rather than a plain 404, because 410 tells a search engine the page is permanently gone and gets it dropped from the index faster than 404's "temporarily missing."

And one trap I actually fell into: the redirect rules were written correctly, but the edge router only runs the Worker for explicitly registered paths — a rule without a registered path is dead code. Correct code, passing tests, green build, and production still returns 404. The lesson generalizes to every platform:

After a redirect ships, test the status code with curl -I against the real domain. Passing locally is not evidence that it works in production.

The boundary of this example: it shows that the process for a single-site move and a retraction works. It is not evidence of multi-site governance outcomes. I do not have GSC and CRM data for a dozen sites, so everything in this post about multi-site work is a method framework, not a measured result.

Action for this section: annotate every 301 in your config with its creation date, and put an annual redirect review on the calendar. The default action is keep; deleting one requires a written reason.


Cross-domain competition: differentiate, freeze, or merge

Cross-domain is harder than on-site, because it involves people and their performance targets, not just pages. [FRAMEWORK] Work strictly through four levels and do not skip:

Level 1: assign ownership. Cheapest, lowest risk. Write down "site A owns heavy-duty pallets for Europe, site B owns hygienic pallets for North America" and most conflicts evaporate without a line of code.

Level 2: reposition the secondary site. It stops competing head-on for the cluster and moves to the angle it is uniquely qualified for — localized service, a specific certification, a specific industry. It may still publish about hygienic pallets; it just no longer runs a procurement page competing for the buying intent.

Level 3: freeze. A freeze has to be defined before it can be executed: stop creating new same-intent procurement pages on the weaker site; do not delete, noindex, or 301 the existing ones; stop investing further optimization in them (no keyword expansion, no new internal links); observe for 8–12 weeks to see whether the primary site absorbs the traffic and inquiries. A freeze is reversible, which is its main value — at the end of the window you choose between rolling back (differentiate again) and going forward (merge).

Level 4: cross-domain 301 or full merge. Only with sufficient evidence: two consecutive quarters with no qualified inquiries for that cluster on the weaker site, no independent backlink value, and a primary site that has demonstrably absorbed the demand.

Why not solve it with a cross-domain canonical

A canonical tag tells Google that the official version of this page lives at another URL. But —

[OFFICIAL] canonical is a suggestion, not a command. Google decides which URL best represents the group. The published signal strength ordering is: 301 redirect (strongest) > rel=canonical (strong) > sitemap declaration (weak).

Cross-domain canonicals routinely fail between sites you own: you declare one, but the two pages differ enough that Google decides they are not duplicates and ignores the declaration — so you believe the conflict is resolved while the two pages keep competing, with no error message anywhere. [FRAMEWORK] Consolidate across domains with a 301 (when content and intent really are the same) or with differentiation (when they are not). Cross-domain canonical is genuinely useful in exactly one scenario: a partner legitimately republishes your content and you are declaring original authorship.

Tool limits to understand before a full merge

[OFFICIAL] The Change of address tool has three hard limits: domain-level moves only, so it cannot handle a path-level migration like example.com/petstore/; you must own and have verified both properties; and it does not carry subdomains (including www) — each one is a separate request.

So "move only site A's pallet line into site B" cannot use the tool at all. That job is per-URL 301s, sitemap and internal link updates, and waiting for a recrawl. Underestimating that workload is a common reason cross-domain merges go wrong.

Action for this section: complete level 1 (ownership) and level 3 (freeze), and observe for a full quarter. Any proposal that skips the first three levels and opens with a cross-domain merge should be sent back.


Can a central keyword table survive contact with reality

It can — but only at a fraction of the size most people picture.

The version that fails looks like this: every long-tail query typed into a spreadsheet, each assigned an owning site, every other site "forbidden" from touching it. The failure modes are specific: long-tail queries are infinite and manual maintenance never keeps up; "forbidden" has no exception path, so a real business need routes around it; and once it has been routed around one time, nobody takes the table seriously again.

A workable MVP structure

[FRAMEWORK] Register high-value clusters only:

Cluster IDDemand / intentMarket & languagePrimary domain & pageWhat other sites may doOwnerStatusLast review
PAL-HYG-001Hygienic pallet procurementEN / US, UKSite B /hygienic-plastic-pallet/Food-industry cases and cleaning guides allowed; no new same-intent procurement pageZhangactive2026-07
PAL-HD-002Heavy-duty pallet procurementEN / EUSite A /heavy-duty-pallet/Warehouse automation applications allowed; no new procurement pageLiactive2026-07
PAL-HYG-003How to clean palletsEN / globalSite B /blog/clean-plastic-pallets/Any site may cite and link; no same-title pageZhangwatching2026-06

Note the third row: informational clusters need registering too, precisely because they are the type every site duplicates.

Eight rules that keep it alive

[FRAMEWORK]

  1. Register the top 20–100 high-value clusters, not every query. The long tail is governed by principle, not by rows.
  2. Every new product or solution page must carry a Cluster ID. Not being able to supply one means the positioning is unclear — that interception is itself valuable.
  3. Blog content may cover long-tail questions but may not copy the primary page's procurement intent. Make the test hard: any page whose title contains supplier / manufacturer / for sale / price goes through registration.
  4. Replace "forbidden" with "permitted scope + exceptions," and provide a route to request an exception. Rules that are too rigid simply get bypassed.
  5. Review two things monthly: what shipped this month, and whether the top 20 clusters have new conflicts. No full audits.
  6. One person must hold decision rights. A coordinator without authority is not an owner, they are a note-taker.
  7. Every row carries a "last review" date; anything older than six months drops to stale and is forced onto the monthly meeting agenda.
  8. If performance is measured per site, add a shared global metric. This is the most-skipped rule and the most fatal: if conceding a cluster lowers someone's numbers, no rule survives the pressure.

The publishing gate that gives the table teeth

The table has no authority; the gate does. [FRAMEWORK] The minimum viable version: add a required clusterId field to the content workflow, with three automated checks before publishing:

1. Does clusterId exist in the register?          → otherwise reject
2. Is that cluster's primary domain = this site?  → otherwise require the "permitted scope" justification
3. Does the title hit procurement-intent terms
   (supplier / manufacturer / price / for sale / buy)
   while this site is not the primary domain?     → route to manual approval

The three checks are under a day of development, and they move governance from "relying on goodwill" to "relying on process." A keyword table without a gate typically survives three months.

Action for this section: build the table with 10 rows. Once you can maintain 10, grow to 50. A 500-row table on day one maintains zero rows.


Next: who owns it, a 30-day plan, and the templates

The judgment framework ends here. Turning it into a schedule and a set of files is separate work, which I split into a companion toolkit:

Multi-site SEO governance toolkit: a 30-day plan and three templates

It covers the three things this post does not: a federated split of responsibilities for a team of 7–8 (who decides, and why performance measurement must be dual-track), a 30-day rollout plan (weekly deliverables and exit criteria), and three copy-ready tables — the cluster register, the 301 merge checklist, and the cross-site conflict log.


Conclusion: governance is not eliminating overlap

Mature multi-site governance does not aim for zero keyword overlap between a dozen sites. That is neither achievable nor necessary. The real goal:

Every high-value search demand has a designated primary page, a named owner, verifiable evidence, and an executable exit path.

The primary page answers "which page on which domain serves this demand." The owner answers "who is accountable for the result and who resolves conflicts." The evidence is landing-page inquiries and sales feedback, not ranking screenshots. The exit path answers "under what conditions do we stop investing in the weaker site, merge, or roll back."

Keyword overlap is not the disease. Nobody knowing who is responsible, and nobody being able to prove which page produced the business — that is the disease.


The boundaries of this post

Three kinds of content, three levels of reliability:

What Google says (tagged [OFFICIAL]): written in Google's own documentation, each with a source link. Cite it freely.

What I derived (tagged [FRAMEWORK]): how the scorecard is weighted, how much SERP overlap makes one cluster, what order to work cross-domain conflicts. These come from the official rules plus delivery experience. Use them, but if your data says another approach works better, follow your data — a framework is a starting point.

What is invented (tagged [EXAMPLE]): company P, its 12 sites, every impression and inquiry number. Fictional, representing no real client, and not a promise of results.

What I do not have: any real company's domain list, internal keyword assignments, headcount, performance system, or per-site GSC/GA4/CRM data — and no "here is how much it grew after we did this" conclusion. I do not hold that data, so none of it appears here. The domain migration in "Why a 301 stays up for at least a year" is my own blog and genuinely happened, but it is one site moving house, not evidence of multi-site governance outcomes.

Comments →

CC BY-NC-SA 4.0

Comments

Comments are powered by GitHub Discussions. Sign in with GitHub to comment. Open the matching Discussion