TL;DR
Google runs site-wide quality classifiers alongside its page-level systems. Websites that choose to scale out thin, batched or unoriginal content and the classifier can suppress the whole domain, including your good pages. It mostly gets re-evaluated at core updates, this meaning recovery is slow, and the fastest fix is usually removing weak content, not polishing it.
You can write a genuinely good article, on a site that Google has decided it does not trust, and that article will not rank.
Not because of anything on the page or anything to do with the article, but because of everything around it.
Google does not just score pages, it scores sites.
And once a site-level classifier turns against you, page-level fixes stop working, which is why so many post-update recovery efforts fail.
We know because we watched it happen to one of our own publications, and we are going to show you how this works and why this happens.
What a site-level classifier actually is
A classifier is a machine-learning system that puts things in buckets.
Page-level systems ask “is this page helpful for this query?” A site-level classifier asks a colder question: “is this site, as a whole, the kind of site we want to send people to?”
The answer becomes a sitewide signal that weighs on everything you publish.
Good page on a trusted site: it gets its shot.
The same page on a site the classifier has marked down: it starts the race with a handbrake on.
That is the part page-by-page SEO audits completely miss.
The evidence site-level scores exist
Google is coy about this being the mechanism to how site level scores work or even exist, so it is worth laying out the receipts.
Four independent lines of evidence point the same way.
1. Google’s own documentation
When the helpful content system launched, Google described it explicitly as a sitewide classifier: a signal that applies across the whole site, and one that can take months to lift after you fix things. That system has since been folded into the core algorithm, the sitewide behaviour did not go away, it just became less visible.
2. The Content Warehouse leak
The 2024 leak of Google’s internal API documentation revealed attributes like siteAuthority and site-level quality scores attached to domains, not pages. Whatever Google says publicly about “domain authority”, its own systems store site-level quality measures.
3. The DOJ antitrust trial
Court documents surfaced Q*, described as a site-level quality signal used in ranking. Under oath beats blog posts.
4. Observed behaviour
The pattern every SEO has seen since 2023: sites that scaled thin content losing traffic across every page at once, on the same day, in the same update. Individual pages do not all fail simultaneously. Classifiers flip simultaneously.
What it looks like when one flips on you
This is where we show our own homework.
One of our own publications ran a content-scaling play: 50+ articles published in a matter of days, templated formats, broad topics beyond its core focus.
The kind of thing half the industry was doing.
It worked, at first. The site grew rapidly, peaking around 80000 daily impressions and 100+ daily clicks. Then a spam update landed in Aug 2025, and it lost effectively all of it.
Not a slow decline.
A cliff, across every page, together, the fingerprint of a site-level classification, not a collection of pages underperforming.
The tell is the synchronisation.
When 75 pages lose visibility on the same date, the problem is not 75 separate pages.
It is one decision Google made about the domain… No manual action recorded, just 0 visibility.
Why we are telling you this: Most agencies only publish their wins.
But this is the single most useful dataset we own on how site classifiers behave, and it is exactly the experience that informs how we protect client sites.
We would rather show you the scar and the lesson than pretend we have never seen the inside of one of these.
What trips the classifier
Scaled, templated content.
Many pages, one pattern, minimal unique value per page.
This is precisely what the June 2026 spam update formalised as scaled content abuse, and we are still seeing strong impacts against batched / scaled content from this update.
Topic drift – Publishing far outside your core subject to chase volume. It dilutes what the site is about, and the classifier reads the whole corpus.
A corpus that skews thin.
The ratio matters.
A few weak pages on a strong site are noise.
When weak pages become the majority, they become the site’s identity.
No evidence of experience – Sites with no authors, no first-hand insight, no original data give the classifier nothing to trust.
Why improving pages does not fix it (and what does)
Here is the counterintuitive part.
Once a site-level classifier has turned, rehabilitating your weak pages one by one is usually the trap and incredibly time intensive.
The classifier does not grade your effort, it grades what your corpus is.
This means that polishing 60 thin pages still leaves you with 60 thin-ish pages defining your site – doesn’t really correct the problem.
The faster path is to change what the site is, and the bluntest lever is deletion.
Remove the off-topic drift entirely, merge the cannibals into single strong pages, and keep only what you would rebuild with real data or real expertise.
We call it deleting the denominator: every weak page you remove instantly raises the average quality of everything left.
Then rebuild the identity with the content only you can make, original data, first-hand experience, genuine depth, and make sure your internal linking points at it.
You are not trying to sneak past the classifier.
You are trying to genuinely become a different site.
The recovery timeline (be realistic)
Site-level classifications are mostly re-evaluated at core updates, not continuously.
That has two blunt implications. First, there is a deadline: your cleanup needs to be finished before the next update evaluates you, not started.
Second, recovery is measured in update cycles, typically one to two after a genuine cleanup, which can mean months of suppressed traffic even when you have done everything right.
Two things are not gated while you wait: PR and AI visibility.
Links earned through digital PR are among the strongest recovery signals you can send, and AI engines citing your original data does not depend on the suppressed rankings.
Both are worth pushing hard during the recovery window.
The DNHQ take
Site-level classifiers are the reason the scaled-content era is ending, and honestly, we think that is a good thing, even having taken the hit ourselves.
The sites that win from here are the ones whose whole corpus earns trust: original data, real expertise, and nothing published just to fill a keyword gap.
It is exactly how we approach SEO for clients, and it is why we publish our own research rather than volume content. If your traffic fell off a cliff at a core update and page fixes have not moved it, our SEO team can run the corpus-level audit that page-level tools miss.










