- NO.
- 024
- DATE
- Updated 2026-08-10
- READ
- ~18 min
- KIND
- Notes
- STATUS
- Reviewed
Cutting Metrics: Why Is This Number Here?
A method never run on a client: one question that filters vanity metrics, a three-layer report structure, signal-based staging, and a worked example.
This describes a method I have never implemented.
During my time on the agency side, I produced monthly reports from a template: pull from GSC and GA4, arrange the trends, write next month's recommended actions. Presenting to the client was usually the team lead's job. The only thing I controlled was how the data was organized — which numbers went in the body, which went in the appendix.
Back then I kept hitting the same question: what earns this number its place in the body? I never solved it. Whatever the template said, I delivered.
What follows was designed after leaving, and has not been run on any client. I am writing it because the problem is real, and a scheme thought through but unverified is more useful than a vague complaint. Read it on those terms: every judgment below is at the evidence level of "I think," not "I tried."
The statistics are not repeated here. The argument that one impression with zero clicks proves nothing is worked out in multi-site SEO governance.
One filtering rule
For a metric to stay in the body of a monthly report, it must answer this:
If it changed, what does the person reading this report do differently next month?
Answerable, it stays. Unanswerable, it moves back.
This is not asking whether the number is important. Almost any number can be argued into importance. It is asking which action it connects to.
"Organic impressions up 18%" — and so? If the answer is "continue as planned," then the action is identical to what it would be at down 18%. That number produced no decision value this month.
The word "whose" cannot be dropped. An implementer has dozens of possible actions: rewrite titles, add interior pages, adjust internal links, add case studies. A client usually has one dial: keep spending, or stop.
So the same sentence, "impressions up 18%," may connect for the implementer to "which category of terms is rising, add content there next month," and connect to nothing at all for the person paying. Without naming the subject, this rule becomes a convenient weapon: declare anything the client cares about a vanity metric.
How long a window
The rule deliberately does not say "this month." Saying it would be wrong.
Single-period movement is mostly noise. In the worked example below, qualified inquiries are 3, up from 1 last month — a 200% increase. It sounds like a reason to raise budget. Statistically, 1 to 3 is nothing, and falling back to 1 next month would be perfectly normal. What can drive an action is not the jump but whether it holds for three months.
So every metric in the body also needs a stated observation window and a reason for that window. Window length is set by the metric's noise level, not by the report being called monthly.
| Metric | Window | Why |
|---|---|---|
| Qualified inquiries | 3-month rolling | Monthly counts are often single digits; a one-month delta is unreadable |
| Priority keyword positions | 3 months | Daily volatility is high; the monthly average is already smoothing, so read direction |
| Keyword category coverage | 1 month | Direction changes fast in the early stage; a month carries information |
| Indexing status | 6 months | Changes slowly; monthly views cannot show structural problems |
Trend charts are the easiest thing in a report to lie with. Pick a trough as the starting point and any metric becomes a rise. So a trend must carry two things: a comparison baseline, and how the start point was chosen. Both live in the appendix; the body needs one line — "start point 2026-01, rationale in Appendix B."
A vanity metric is not a useless number
Treating vanity metrics as a fixed blacklist collapses on contact with real projects. Impressions, average position, indexed count, bounce rate — none is inherently worthless.
A site launched three months ago has zero inquiries and empty landing-page data. Impressions are the only moving signal. They show Google beginning to understand what the site is about, and which category of terms it recognized first. Cut that and you have cut the entire report.
The same impression count on a two-year-old site with steady monthly inquiries is a placebo. That site should be reading which landing pages produce inquiries and which pages are declining.
A vanity metric is not a useless number. It is a number with no corresponding decision.
The same metric can be the key indicator on one site and decoration on another. So cutting metrics is not writing a blacklist; it starts with asking which stage this site is in.
A report is not one thing; it is three layers
Treating a monthly report as one homogeneous document is a mistake. It has at least three layers, with different readers and different acceptance criteria:
| Layer | Answers | Read by | Filter |
|---|---|---|---|
| Delivery | What we did this month | The person paying | The rule does not apply; this is contractual evidence |
| Diagnostic | What happened on the site, and why | Implementer and decision-maker | "If it changed, whose action differs?" |
| Decision | What to do next month, and why this | Decision-maker | Every item must land on a concrete action |
That split answers one objection. A client will say: we published 8 articles and fixed 3 technical issues — my actions do not change either way, so by your rule all of that should be cut?
No. A delivery list is not a metric; it is proof of performance, and measuring it with the diagnostic ruler is the wrong instrument. Conversely, an honest delivery layer relieves pressure on the diagnostic layer: state plainly what was done, and you no longer need a screenful of numbers to imply "we were busy."
Cutting metrics requires a complete appendix
From the client's seat, cutting metrics is reducing disclosure. There is no getting around that.
An agency that wants to obscure things can use this rule to cut every unfavorable number and sound principled doing it. "Cannot answer an action" is easy to fake. The method does not prevent bad faith.
Its only defense: a cut metric is not deleted, it is moved to the appendix — and the appendix must be complete and checkable.
The report is a decision document; the reader finishes it and decides what to do next month. The appendix is evidentiary material; its reader is someone three months later who wants to verify what was actually true at the time. Different readers, different purposes.
Whether a report has earned the right to cut metrics rests on one test: can any cut number the client wants to look up be found in the appendix, in its original definition?
If not, the body should keep carrying everything. When the appendix is incomplete, cutting is not information design. It is reduced disclosure.
A staged metric table, switched by signal
| Site stage | What belongs in the body | What moves back |
|---|---|---|
| New / zero baseline | Which keyword categories have impression coverage (not the absolute number), how many target pages are indexed, first terms entering the top 20 | Inquiries, conversion rate, CTR trends. Sample too small; you could only write "cannot yet judge" |
| Ramping | Inquiries by landing page, terms entering the top 10 and their landing pages, which pages concentrate clicks | Site-wide average position, total impressions. Totals begin masking structural change |
| Steady | Change in inquiry-source pages, priority pages dropping out of the top 10, competitor movement on priority terms | Indexed count, impressions, average position — unless one is anomalous enough to warrant investigation |
Stages are not defined by month count. "Months 0–6 count as a new site" does not survive one question: what is a site with zero inquiries in month 7?
Define them by signal instead:
- The first qualified inquiry attributable to a landing page → ramping.
- Qualified inquiries non-zero for three consecutive months → steady.
Write the downgrade too, and treat it as more important than the upgrade. If a steady-stage site has zero qualified inquiries for three consecutive months, it returns to the ramping column. When a site is in trouble is exactly when the report most needs to switch back to structural metrics. A table that only defines upgrades is empty at the moment it matters most.
Seasonal industries are an exception and need separate handling. Christmas goods, agricultural machinery: zero inquiries in the off-season is natural, and the rules above would toggle the stage twice a year. For those, the switching signal compares to the same period last year, not the previous period: three consecutive months that had inquiries last year and have none now.
The "new site" cell reads categories, not volume. AI overviews and zero-click results keep decoupling impressions from clicks, so "impressions are up" is itself depreciating. But the question for a new site was never "how many people saw it." It is "has Google begun to understand what this site does?" If the categories are right, add content there next month; if they are all irrelevant long tail, the positioning is not landing, and next month's change is positioning, not production volume. Even with clicks fully decoupled, it still connects to an action.
The right-hand column is not deleted. It is moved back.
Metric inflation is defensive writing
The usual explanation for metric inflation is "clients like these numbers." That absolves the writer too cleanly. The more common cause sits with whoever writes the report: adding one more number is safer than writing one judgment.
A wrong number can be defended with "that is what the data said." A wrong judgment is "you judged it wrong." A report stuffed with twenty metrics contains a sentence corresponding to any outcome; a report with five metrics, each followed by "so next month we do X," can be audited a month later.
But framing this as courage is wrong. When a judgment turns out wrong, the person taking it in front of the client is usually not the one who wrote the report — it is the one who presents it. When the writer of a judgment and the bearer of its consequences are different people, more numbers and fewer positions is the rational response.
Metric inflation is not a character problem. It is a risk-distribution problem.
That diagnosis yields something actionable; "be brave" does not. Since the problem is risk distribution, label each judgment's confidence and deliver it together with its risk:
- [Confirmed] — directly supported by data, safe to state externally.
- [Inferred] — my most likely explanation, alternatives not excluded, to be tested next month.
- [Unknown] — I do not know; here is what we did this month to make it knowable.
An [Inferred] judgment can be read out verbatim by the presenter without underwriting it. If it is disproved a month later, the reconciliation is "this was labeled inferred," not "you lied to me last month."
Labels only count once they close the loop
All three labels look forward. Written and never settled, they are just a disclaimer.
So the first item in each report is settling the previous report's [Inferred] and [Unknown] items. Each gets an outcome: confirmed, disproved, or still unknown. Disproved items state the new explanation.
Without that step, a year produces a dozen inferences floating in old documents and nobody knows how many were right. With it, the report starts having a memory. A report with a memory is what lets a client judge whether this team is reliable.
But the client wants ranking screenshots
I have no clever counter, only two moves:
- Do not argue; change the structure. Include the screenshot, in the appendix, and give the first screen of the body to inquiries and landing pages. Arguing that a metric is meaningless is arguing an abstract proposition and rarely wins; placing two kinds of number in two places expresses priority through layout.
- Attach an action to the ranking. "Priority term A fell from 12 to 19; its landing page has not been updated in three months; next month we add the certification material and a specification table." Now it connects to an action and is no longer a vanity metric.
Move 1 has a side effect: a silent demotion reads as guilt. Quietly relocating the screenshots makes a client think "what is he hiding," not "he is improving the information structure." So state the reason in the body when you move it — one line is enough: "Ranking detail is in Appendix A; this screen is reserved for inquiry sources, because ranking is process and inquiries are outcome."
What a report written this way looks like
All data below is a fabricated example and comes from no real client. Setup: a B2B industrial equipment export site, live for 8 months, first attributable inquiry last month, just entered the ramping stage.
Settlement of the previous period
- Last month's [Inferred] "the homepage redesign drove brand-term click growth" is disproved this month. Brand clicks began rising two weeks before the redesign; the timing does not line up, and the actual cause was an industry trade show. The redesign's effect remains unknown.
- Last month's [Unknown] "whether category pages are fully indexed" is still unknown. The indexing report lags; it stays open.
Delivery layer | what we did
- 6 new product pages (XX-200 series) and 2 case pages.
- Fixed self-referencing canonical errors on paginated category pages, affecting 41 URLs.
- Updated specification tables on 3 older product pages.
Diagnostic layer | what happened on the site
- [Confirmed] 3 qualified inquiries, all landing on
/products/xx-200/and/cases/food-line/, versus 1 last month. The 3-month rolling series is 1 / 1 / 3 — too small to read as a trend, so this month records it without interpreting it. - [Confirmed] Terms in the top 10 rose from 4 to 11, of which 9 are "model + specification" queries matching this month's new pages. Window is 1 month, because category change is fast enough in the ramping stage.
- [Inferred] Inquiries concentrate on product pages rather than case pages. Two factors — the specification table and the form position — are currently entangled, so which one is working is unclear. Next month we change only the form position and leave the tables alone, to separate them. Even so, 3 inquiries cannot settle it; a real read arrives in month 3.
- [Unknown] The canonical fix has no visible effect yet; indexing status lags. Re-check those 41 URLs next month.
Decision layer | what we do next month
- Build XX-300 and XX-500 series pages on the XX-200 structure, 4 pages each. Basis: 9 of this month's new top-10 terms follow that pattern.
- Move the case-page form above the fold, leaving specification tables unchanged, to separate the two factors in the [Inferred] item above.
- Re-check indexing on the 41 URLs; if unrecovered, investigate server-side rendering.
Appendix
- Raw exports:
exports/2026-07/gsc-query-page.csv,ga4-landing.csv. - Data definition: GSC property
sc-domain:example.com, 2026-07-01 to 07-31, all countries, search type web. - Pull date: 2026-08-03, after the data delay window.
- Trend start 2026-01, chosen as the first complete calendar month after launch.
- Detail of the 11 top-10 terms: page 2 of the csv.
The body carries no total impressions, no average position, no indexed count. Not because they are unimportant, but because on a ramping-stage site, none of the three decisions above changes whether those numbers move. They are in the appendix csv; they have not disappeared.
For a blank version you can fill in directly, this site's library has the GSC / GA4 SEO monthly review template, with a Markdown report and a CSV action ledger.
Template v1.1 has absorbed this method — the three-layer split, observation windows, confidence labels with previous-period settlement, and signal-based staging — and every one of those blocks is marked "unverified" in the template, because that is their evidence level. Template 1.0's data definitions, four-question framework, and pre-publish QA are unaffected; those are a different matter.
How to start in month one
You do not need to convince anyone to change the template first. All three steps sit within one implementer's own authority:
- Do not change the report; run one filter pass. Open last month's report and ask of every number, "if it changed, whose action differs?" Cross out the unanswerable ones. The output is not a new report; it is a list you now hold in your head.
- Build the appendix. Add one section at the end of the existing template containing four things: where the raw exports live, the data definition, the pull date, and the trend start point with its rationale. This changes nothing anyone sees in the body, and it is the precondition for every later relocation. Without an appendix, moving equals deleting.
- Try confidence labels in the decision layer only, and settle them at the top of the next report. Prefix next-month recommendations with [Confirmed]/[Inferred]/[Unknown], then write the outcome of each at the start of the following month. Labels and settlement ship together; labels alone are only a disclaimer.
The parts that genuinely need a template change, like the three-layer split, come last — after the first three steps have produced a month or two of real experience.
Four places this design collapses
I have solutions to none of these, only partial responses.
One: it may be filtering my imagination rather than bad metrics.
"Cut what cannot answer an action" hides a premise: that being unable to answer proves the number connects to no action. More often, it only proves I do not know what to do.
Concretely: "mobile CTR is 40% below desktop and has been for three months." Someone who knows the domain immediately checks mobile title truncation, above-the-fold structure, and load speed. Someone who does not cannot name an action, and by the rule moves it to the appendix — where it may have been the single most worth investigating row in the table.
So this rule's resolution ceiling equals the experience ceiling of whoever applies it. It filters out "metrics the person in this seat cannot think of a move for," not "metrics without value."
Two partial responses. First, before cutting, ask one more question: does it truly connect to no action, or do I not know what to do? Record the latter in the appendix as [Unknown] rather than discarding it. Second, have someone else ask the question — which is point four.
Two: a client in the investment phase sees no money, so why renew?
The hardest one. SEO's first months often produce no inquiries, and by the staging table the body contains only keyword categories and indexing. Those connect to none of the paying person's actions; their dial is renew or not. Measured by this rule, the rule exposes its own hole.
Partial response: what connects to a client action during the investment phase is not any single-period number, but the reconciliation of last period's promise against this period's delivery. "Last month we said we would get category A terms indexed; this month we have 37, of which 29 are in the target category" — the client's action is judging whether this team does what it says. That is exactly what the "labels only count once they close the loop" section is for. But three consecutive months of unmet promises cannot be saved by any layout.
Three: making inquiries the gold standard invites junk inquiries.
Measured on count alone, the cheapest path is chasing "free," "cheap," "DIY" queries. The number looks good and sales receives students and competitors. That is a perverse incentive created by the metric itself.
Partial response: inquiries must carry a qualifier from the sales side, at minimum "qualified / unqualified," with the determination belonging to sales, not SEO. Multi-site SEO governance uses the same principle for commercial-value scoring. The unanswered part: many projects simply cannot obtain sales-side feedback. I have no substitute there, only the obligation to state in the report that the metric is crippled.
Four: it will degrade into a new dogma.
"Every metric must be followed by an action" degrades easily into a forced platitude — "impressions rose, we will continue" — formally compliant and substantively empty.
The only defense is that the question must be asked by someone else. You can always invent an action for your own numbers; asked directly, "so which page are you actually changing next month," the ones who cannot answer become visible immediately. So this only works in a team where someone reviews.
Boundaries
- This method has not been implemented on any client project, nor tested through a complete report cycle. It is a design, not a retrospective. The example report is fabricated and contains no real client's identity, data, or screenshots.
- The staging signals and observation window lengths are my own values, not an industry standard. Site scale and competitive intensity will stretch or compress them.
- No official documentation is cited here. The official basis for GSC data definitions is sourced item by item in multi-site SEO governance, and the division-of-responsibility framing is in the B2B delivery framework.
Comments
Comments are powered by GitHub Discussions. Sign in with GitHub to comment. Open the matching Discussion