- NO.
- 005
- DATE
- READ
- ~7 min
- KIND
- Case study
- STATUS
- Reviewed
5,794 Pages, 254 in the Sitemap: a Teardown
A technical SEO teardown of a B2B export site: the sitemap declares 254 URLs while search engines know 5,794. The problem lives outside the sitemap.
This is a de-identified rewrite of a technical SEO diagnosis. The original subject is a publicly reachable B2B site. No client identity, no private screenshots, and no claims about downstream business results — only the reviewable diagnostic path, public page signals, and de-identified metrics.
The question that started it:
Why does Google appear to know about 5,794 pages?
My conclusion: this site does not have too few pages. It has the valuable ones under-declared and the worthless ones taking up room. The sitemap is disciplined — 254 URLs. The problem is everything outside it, where search engines discovered a mass of tag archives, date archives, and leftover pages on their own.
Background
A B2B export site of this kind usually has three goals:
- Let buyers quickly understand the products, the capability, and whether the company is credible.
- Let search engines reliably find the high-value pages instead of getting lost in archives, parameters, and history.
- Keep content expansion controllable, so the site does not become a pile of low-value URLs in three years.
This was a first-round technical teardown, not full-service operation and not deep GSC attribution. The job was to determine what the number 5,794 actually represents.
The problem
The sitemap declares 254 high-value URLs. Search engines know about roughly 5,794 — about 22× more.
That gap is not sitemap bloat. It is a large population of indexable, auto-generated, low-value pages living outside the sitemap. They may not exhaust crawl budget immediately, but they dilute the site's quality signal, scatter internal link equity, and make index hygiene look poor.
So the right question is not "why does the sitemap have so many pages?" It is:
Where did the 5,500-plus URLs outside the sitemap come from, and which of them should leave the index?
Ask it that way and the fix direction follows.
Quantifiable findings
This round uses public pages and de-identified numbers only:
- Total declared in sitemaps: 254 URLs.
- Pages known to search engines: ~5,794.
- Gap: about 22×.
- Homepage HTML weight, sampled: ~157KB.
- Category page HTML weight, sampled: ~218KB.
- Date archives: sampled paths return 200, no
noindexfound. - Tag archives: sampled paths return 200, no
noindexfound. - Author archives and site search: already carry
noindex— so not every archive is out of control.
These numbers are enough to set direction. They are not enough to claim an outcome: after the fixes ship, it takes 4–8 weeks of GSC observation to say what changed.
The actions
A first round does not require buying a stack of tools. It can be this plain:
- Fetch
robots.txt, the sitemap index, and every child sitemap. - Count declared URLs per sitemap.
- Fetch the homepage HTML and read title, description, canonical, H1, and template leftovers.
- Probe the common WordPress bloat URLs: tag archives, date archives, author archives, search pages.
- Record the HTTP status of each URL type and whether the page carries
noindex.
That establishes direction. Exact composition still needs the GSC "Pages" report plus a Screaming Frog (or equivalent) crawl to confirm.
Output
The deliverable here is not a traffic report. It is a prioritized fix list:
- P0: handle index bloat outside the sitemap.
- P0: clear the leftover template metadata on the homepage.
- P1: fix the hidden, meaningless H1.
- P1: reconcile and merge the mixed content types.
- P2: keep verifying the history gap and page weight.
P0: index bloat is the actual problem
The evidence is direct:
- Sitemaps declare 254 URLs in total.
- The date archive
/2025/returns 200 and is indexable. - Tag archives at
/blog_tag/.../return 200 and are indexable. - Author archives (
?author=1) and site search (?s=...) are correctlynoindex.
Worse, some tags are placeholder fakes like blog-tag-1, blog-tag-2. They carry no editorial value and are being discovered and evaluated anyway.
The usual shorthand for this is "wasted crawl budget." For a small-to-mid B2B site, crawl budget is probably not the first-order problem. The practical risks are:
- A few thousand thin pages dilute the site's overall quality signal.
- Internal link equity flows to archive pages with no commercial value.
- The structure a search engine sees gets dirty, so the pages that matter — products, services, cases — stand out less.
The fix:
- Set tag and date archives to
noindex, follow. - Delete or merge the
blog-tag-1/2style fake tags. - Use GSC to break the total into indexed / crawled-not-indexed / duplicate / alternate page / discovered-not-indexed.
- Resubmit the sitemap after cleanup and watch indexed pages and crawl stats for 4–8 weeks.
P0: leftover template metadata on the homepage
The homepage <head> still carried metadata from the purchased HTML template — the title and description of a car dealership demo.
It looks like a trivial bug, and it damages credibility disproportionately. On a marine hardware or industrial site, dealership demo copy in the head produces three problems:
- Title and description send conflicting signals.
- It advertises that the site is a hard-modified template.
- It reads as untrustworthy to buyers and search engines alike.
The fix is small: remove the demo metadata from the theme header template, keep the real title, description, canonical, and Open Graph data.
P1: a hidden, meaningless H1
The homepage carries this:
<h1 style="display:none;">Home</h1>
The H1 is one of the page's topic signals. Writing it as Home and then hiding it with display:none wastes the single most prominent heading slot on the site.
Better: a visible, natural H1 containing the core business term. For example:
Marine & Boat Hardware Manufacturer
The actual wording follows the real business, the primary market, and the conversion goal. Do not stuff keywords into a heading no human would read.
P1: two content types living side by side
The sitemaps show two article types coexisting: a large number of post entries and a small number of blog entries, with discontinuous year coverage.
That usually means one of:
- A theme or plugin migration left the old content type behind.
- The editorial team used
postfor one period andblogfor another. - Some content was bulk-generated or imported.
You cannot conclude "low quality" from the structure alone, but you must sample it — especially the heavy 2025–2026 clusters. Check whether those pages are original, whether they serve a buyer, and whether they are synonym-swapped AI bulk pages.
The fix is to pick the canonical type first, then merge, redirect, or noindex the redundant one. Do not let two article systems coexist indefinitely.
P2: history gap and page weight
Two secondary findings.
First, a history gap: the blog sitemap covers several years but one year is missing. Confirm whether content was lost in migration, never published, or is now 404 with no 301.
Second, page weight: ~157KB of homepage HTML and ~218KB on category pages. That does not by itself equal a Core Web Vitals problem, but it suggests a heavy template and is worth verifying with PageSpeed Insights or CrUX.
What the site already got right
A diagnosis that only lists faults is not a diagnosis. This site has fundamentals worth keeping:
- The homepage title and description carry clear keywords, a selling point, and intent.
- The canonical is self-referencing.
<html lang="en">is correct.- Author archives and site search already carry
noindex. - The sitemap uses an index structure and
lastmodvalues are current.
So the fix is cleanup of index hygiene and template residue, not a rebuild.
What transfers
The reusable procedure:
- Separate "URLs you submitted" from "URLs search engines discovered."
- Count by page type, not by total page count.
- Pull auto archives, search pages, author pages, parameter URLs, and historical leftovers into their own bucket first.
- Set direction from public crawl results, then confirm scale in GSC.
- Attach every recommendation to one shipping action and one follow-up metric.
The first five steps if I took it over
- Pull the GSC "Pages" report and split 5,794 into indexed, crawled-not-indexed, duplicate, alternate page, and discovered-not-indexed.
- Batch tag and date archives to
noindex, follow; delete the fake tags. - Clear the homepage template residue and fix the H1.
- Sample content quality in the high-volume years, especially duplication across product, service, and article pages.
- Resubmit the sitemap and watch index count, crawl volume, and core keyword pages for 4–8 weeks.
The value of a technical teardown is not pasting a tool's error list. It is stringing the problems into one prioritized path someone can actually ship. Here the load-bearing judgment is a single sentence: 5,794 is not a sitemap number, it is an index hygiene problem outside the sitemap.
Comments
Comments are powered by GitHub Discussions. Sign in with GitHub to comment. Open the matching Discussion