Lightscore

Why are my Lighthouse scores different?

Two different problems arrive as the same question. Separate them before you look for a cause.

Noise is a spread. The same tool on the same machine returns 91, then 94, then 88. Nothing about the page changed.

Bias is an offset. Your laptop reports 96 and PageSpeed Insights reports 61, every time you look. Some setting differs.

Noise disappears when you take the median of several runs. Bias never averages out, so you have to find the setting that differs and match it.

A third case fits neither word. If you measure from two countries, the page itself can differ. We measured one article that loads a geo-targeted banner in Brazil and not in the United States. The score moved 38 points while network time stayed the same. That isn't noise, and no setting on your machine causes it — see why your site scores differently in another country.

How do I tell them apart?

Run the same tool on the same machine five times. Write down the five scores. Then compare that spread against the gap that worries you.

yes

no

yes

no

Run one tool
5 times, one machine

Wide spread
across the 5?

Noise.
Take the median.

Other tool outside
that spread?

Bias.
Find the setting.

They agree.
Nothing to fix.

Two questions, one procedure. Measure your own spread before you compare two tools.

Most troubleshooting skips this step. People compare one local run against one PageSpeed Insights run, then spend an afternoon hunting for a cause that doesn't exist.

How much noise is normal?

We keep three example reports warm, so we have a fixed rig and a log of it. Between 1 July and 13 August 2026 that rig produced 130 mobile Lighthouse runs, on one Lighthouse version and one machine class, from two regions. Scores below are the median, with the full range in brackets.

Page Runs Performance Median Total Blocking TimeHow long the main thread was busy and couldn’t respond to taps or clicks.
lightscore.dev 47 100 (99–100) 5 ms
wikipedia.org 44 92 (78–97) 240 ms
vercel.com 39 32 (30–40) 2,846 ms

Across all 130 runs, 80 percent landed within 2 points of their own page's median. 96 percent landed within 5 points. The worst single run sat 14 points away.

So a 3-point drop between two runs isn't a regression. A 14-point drop can still be noise, if the page does enough work for the numbers to wander that far. Our own homepage does almost none, and it returned 99 or 100 on all 47 of its runs.

The score is steadier than the metric under it

This part surprised us.

Wikipedia's Total Blocking Time ranged from 16 ms to 688 ms, a factor of 43, while its performance score moved only from 78 to 97. Vercel's ranged from 1,459 ms to 16,079 ms for a score between 30 and 40.

Lighthouse maps each metric onto a log-normal curve, which is steep in the middle and flat at both ends. That is how Lighthouse turns Total Blocking Time into a score. Vercel sits far past the flat end:

Metric score Total Blocking Time (ms) 0 20 40 60 80 100 0 3,000 6,000 9,000 12,000 15,000 18,000
Lighthouse maps Total Blocking Time onto a log-normal curve. Past about 2,000 ms it is nearly flat, so a 13-second swing is worth three points.

A 13-second swing is worth 3 metric points. That metric carries 30 percent of the performance score, so the whole swing costs under one point overall.

A steady score doesn't prove a steady page. We think the composite is the least useful number in a Lighthouse report — see what a good Lighthouse score actually tells you. Read the five metrics underneath it instead.

what’s this

The small set of field metrics Google treats as a ranking input and reports in Search Console. We measure their lab equivalents here — real-user field data (CrUX) is a separate signal. TBT is the lab stand-in for INP, not a Core Web Vital itself.

What actually causes bias

  1. CPU speed. Lighthouse runs a small benchmark and stores it as environment.benchmarkIndex in the report JSON. Google's scale puts 1000 and above at desktop class. Our fleet returned 1,379 to 2,103 on 129 of those 130 runs, and 7,068 on the last one. Compare this number first.
  2. Throttling method. The default method is simulate. Chrome records the trace at full speed. Lighthouse then models 150 ms round-trip time, 1.6 Mbps and a 4× CPU slowdown. DevTools applies real throttling instead, and the two disagree.
  3. Form factor set without throttling. This one caught us. Pass formFactor: 'desktop' on its own and Lighthouse keeps the mobile connection profile. You then measure a desktop page over a throttled phone connection, and grade it on the stricter desktop curve. Use the desktop config preset, not the flag.
  4. Metric weights and scoring curves change between Lighthouse releases, so two versions aren't comparable.
  5. Latency to your origin and CDNA network of edge servers that serve your content from near the visitor. edge behaviour both change with the region you run from — though less of that latency reaches the score than you would expect.
  6. The page. Google's own notes on Lighthouse variability rate page nondeterminism as a high-impact source. A/B tests, ad slots, cookie banners and Third partyCode and assets loaded from another origin — analytics, fonts, ads, widgets. tags change what loads.

The first four are yours to fix. The last two are measurements, not faults.

What to do about it

  1. Fix one tool and one machine, then run it five times and take the median. Google states that the median of 5 runs is twice as stable as 1 run.
  2. Read benchmarkIndex in both reports before you blame the page.
  3. Record the Lighthouse version beside every score you keep.
  4. Set your pass threshold against a distribution, not one run. Google recommends the median, the 90th percentile, or min and max. If that threshold gates a build, the defaults matter — see how to run Lighthouse in CI.

What we do about it

Every Lightscore run pins the Lighthouse and Chromium version. The runtime line on the homepage names both, and each result repeats the version that produced it. All regions run on one machine class.

The API accepts up to 5 runs per region and returns the median run. The free one-URL tool runs once. That's a single sample, and the table above says what a single sample is worth.

We also hand back every run, not only the median one. checks.lighthouse.per_run carries each run's four category scores next to its benchmark_index. You can measure our noise yourself rather than trust this page about it.

What we can't remove is your page. If your site runs an A/B test, or loads a tag on a timer, our numbers move too. No lab tool fixes that. A tool that promises identical scores every time hides the spread instead of measuring it.

Common questions