Methodology

How we grade the web, in full.

Every grade on our dashboard comes from measurements anyone could repeat — this page explains what we measure, how it becomes a score, and, just as important, what we can't see.Current as of July 2026 · updated as the model evolves.

The short version
Once a week, we visit about a thousand of the most-used U.S. consumer websites with a real browser, behaving like a first-time visitor who touches nothing — no clicks, no cookie-banner choices. We record every network request the site makes. Each site is then scored on two things: how much consent machinery it has built (did it give you a real way to say no?) and how broadly it connects you to third-party tracking companies by default, before you've made any choice. The grade is simply which quadrant of that two-axis map the site lands in: A, B, C, or D.
01
The measurement

A real browser, a first-time U.S. visitor, no interaction

The crawler drives a real Chrome browser from a U.S. vantage point across roughly 1,000 curated consumer sites in 16 industries, weekly. It loads each homepage, waits while scripts fire, scrolls once, and records everything: every network request, the consent banner and its buttons (if one appears), the standard consent interfaces a site exposes to the browser, the cookies that get set, and the policy links in the footer.

It never interacts with a consent banner. Everything we record is therefore the site's default behavior — what happens to a visitor who hasn't said yes or no. In the U.S., where tracking is generally lawful until you opt out, the default state is exactly what most visitors experience.

Failed measurements are discarded, not graded. If a site blocks the crawler or fails to load meaningfully (fewer than 30 requests captured), that's a failed measurement — the site is excluded that week rather than mistaken for a "clean" site.

03
The vertical axis

Tracking restraint — how much does the site hold back, by default?

A score from 0–100 for how much a site holds back from handing your visit to outside tracking companies before you've made any choice. 100 means almost nothing reaches third-party trackers by default; low scores mean many different tracking companies are contacted the moment the page loads.

Behind the score is a count of the distinct third-party tracking companies a site connects you to by default, weighted by how consequential each kind of company is and adjusted for site size. We deliberately count breadth (how many different parties) rather than volume (how many requests): volume mostly measures how chatty a site's engineering is, while breadth measures how widely your visit is distributed.

SeverityWho's countedWeight
Tier 3Identity and audience brokers, retargeters, ad-conversion pixels, real-time bidding — companies whose business is recognizing you across sites4
Tier 2Social-platform widgets and web-to-app identity bridges — your visit phones home to one platform2
Tier 1Operational vendors — ad delivery, affiliates, email tooling, generic personalization, session replay1
Tier 0Infrastructure, first-party services, consent tooling itself0

Behavior outranks labels. If any vendor — whatever its category — is observed performing cross-site identity syncing (cookie-matching endpoints, ID-bridge calls) or hiding behind a disguised first-party subdomain, it's treated as Tier 3 or higher for that site. We verified this rule against a full crawl: it fires almost exclusively on genuine sync endpoints.

Size-adjusted: the severity-weighted count is divided by the square root of the site's total requests, so a large, media-heavy site isn't punished for being big — only for being broad.

Expressed as restraint: the published score is that count inverted onto a 0–100 scale, so more sharing lowers the score and both axes read the same direction. The grade divider sits at exactly 50, and a 100 means no cross-context tracking was observed on our visit — an observation floor, not a certificate that the site never tracks.

04
The grade

The grade is the quadrant — nothing more

Each axis has one threshold, set at the median of every site we measured and held fixed when we revise the model, so grades stay stable week to week. On the published 0–100 scales that threshold sits at exactly 50 on both axes — a site above 50 on both grades A. Which side of each divider a site lands on places it in one of four quadrants, laid out here exactly as they sit on the dashboard:

more tracking restraint →
C · Minimal

Little consent effort, but little tracking either.

A · Restrained

Real consent effort and a light default tracking footprint.

D · Unrestrained

Little consent effort and broad default tracking.

B · Tracks anyway

Real consent effort, yet heavy tracking still fires by default. It's the web's largest quadrant, and our headline finding.

more consent effort →

Every weight and threshold on this page is a setting we publish, not a black box, and revisions are audited against how the standards actually behave in the wild before they ship.

What we can't see

Known limits, stated plainly

One vantage point

We crawl from the U.S. Sites that behave differently for European visitors — or gate tracking only for certain states — show us their U.S. default, nothing else.

Restraint can be invisible

A site that blocks its tags entirely until consent shows us nothing, which looks identical to having no Google tags at all. Our "defaults to denied" signal only rewards restraint we can observe; absence is never counted against a site.

Low-effort sites tie exactly

A site with only a privacy policy scores exactly what every other privacy-policy-only site scores. Those are real ties, and we plot them as ties — we don't scatter them artificially to look more precise than we are.

One visit, weekly

Ad markets are dynamic; the exact vendor set varies visit to visit. Breadth of distinct parties is far more stable than request counts — one reason we score it — but week-to-week movement within a band is normal.

Not everything is classifiable

A small share of tracking calls comes from vendors we haven't yet identified. Unknown vendors count toward breadth only when they're observed doing identity syncing.

Some sharing is off-page

We measure what your browser does. We read what it sends to a site's own collection endpoints and can often name the real destination, but when a site's servers pass your visit along out of the browser's view, those parties don't appear in our count. Our numbers are a floor, not a ceiling.

A fixed observation window

We watch each site for a fixed time from the moment the visit starts, so a fast page and a slow page get equal opportunity to show their trackers. A tracker that loads after the window closes is missed — timing can under-count, but it never invents.

Some sites block measurement

A share of sites put automated visitors behind a security wall, and a walled visit is a failed measurement, not a clean one. We flag those visits rather than mistake a blocked page for a private one.

Revision log

How the model changes over time

Two kinds of change land here. Measurement changes — a new signal, a sharper detector — can move grades. Presentation changes — how a score is displayed — never do. We date and describe both, so a grade that moves is never a silent change.

No revisions logged yet. Entries will appear here as the model evolves.
Observable indicators, not legal conclusions. Grades describe measurable industry practice — what a site's machinery does for a first-time visitor — never legal compliance, intent, or wrongdoing. A grade can change as a site's setup changes, and as this methodology improves.