Why are my Lighthouse scores different?
Two different problems arrive as the same question. Separate them before you look for a cause.
Noise is a spread. The same tool on the same machine returns 91, then 94, then 88. Nothing about the page changed.
Bias is an offset. Your laptop reports 96 and PageSpeed Insights reports 61, every time you look. Some setting differs.
Noise disappears when you take the median of several runs. Bias never averages out, so you have to find the setting that differs and match it.
A third case fits neither word. If you measure from two countries, the page itself can differ. We measured one article that loads a geo-targeted banner in Brazil and not in the United States. The score moved 38 points while network time stayed the same. That isn't noise, and no setting on your machine causes it — see why your site scores differently in another country.
How do I tell them apart?
Run the same tool on the same machine five times. Write down the five scores. Then compare that spread against the gap that worries you.
Most troubleshooting skips this step. People compare one local run against one PageSpeed Insights run, then spend an afternoon hunting for a cause that doesn't exist.
How much noise is normal?
We keep three example reports warm, so we have a fixed rig and a log of it. Between 1 July and 13 August 2026 that rig produced 130 mobile Lighthouse runs, on one Lighthouse version and one machine class, from two regions. Scores below are the median, with the full range in brackets.
| Page | Runs | Performance | Median Total Blocking TimeHow long the main thread was busy and couldn’t respond to taps or clicks. |
|---|---|---|---|
| lightscore.dev | 47 | 100 (99–100) | 5 ms |
| wikipedia.org | 44 | 92 (78–97) | 240 ms |
| vercel.com | 39 | 32 (30–40) | 2,846 ms |
Across all 130 runs, 80 percent landed within 2 points of their own page's median. 96 percent landed within 5 points. The worst single run sat 14 points away.
So a 3-point drop between two runs isn't a regression. A 14-point drop can still be noise, if the page does enough work for the numbers to wander that far. Our own homepage does almost none, and it returned 99 or 100 on all 47 of its runs.
The score is steadier than the metric under it
This part surprised us.
Wikipedia's Total Blocking Time ranged from 16 ms to 688 ms, a factor of 43, while its performance score moved only from 78 to 97. Vercel's ranged from 1,459 ms to 16,079 ms for a score between 30 and 40.
Lighthouse maps each metric onto a log-normal curve, which is steep in the middle and flat at both ends. That is how Lighthouse turns Total Blocking Time into a score. Vercel sits far past the flat end:
- Total Blocking Time of 3,057 ms scores 3 out of 100.
- Total Blocking Time of 16,079 ms scores 0 out of 100.
A 13-second swing is worth 3 metric points. That metric carries 30 percent of the performance score, so the whole swing costs under one point overall.
A steady score doesn't prove a steady page. We think the composite is the least useful number in a Lighthouse report — see what a good Lighthouse score actually tells you. Read the five metrics underneath it instead.
what’s this
The small set of field metrics Google treats as a ranking input and reports in Search Console. We measure their lab equivalents here — real-user field data (CrUX) is a separate signal. TBT is the lab stand-in for INP, not a Core Web Vital itself.
What actually causes bias
- CPU speed. Lighthouse runs a small benchmark and stores it as
environment.benchmarkIndexin the report JSON. Google's scale puts 1000 and above at desktop class. Our fleet returned 1,379 to 2,103 on 129 of those 130 runs, and 7,068 on the last one. Compare this number first. - Throttling method. The default method is
simulate. Chrome records the trace at full speed. Lighthouse then models 150 ms round-trip time, 1.6 Mbps and a 4× CPU slowdown. DevTools applies real throttling instead, and the two disagree. - Form factor set without throttling. This one caught us. Pass
formFactor: 'desktop'on its own and Lighthouse keeps the mobile connection profile. You then measure a desktop page over a throttled phone connection, and grade it on the stricter desktop curve. Use the desktop config preset, not the flag. - Metric weights and scoring curves change between Lighthouse releases, so two versions aren't comparable.
- Latency to your origin and CDNA network of edge servers that serve your content from near the visitor. edge behaviour both change with the region you run from — though less of that latency reaches the score than you would expect.
- The page. Google's own notes on Lighthouse variability rate page nondeterminism as a high-impact source. A/B tests, ad slots, cookie banners and Third partyCode and assets loaded from another origin — analytics, fonts, ads, widgets. tags change what loads.
The first four are yours to fix. The last two are measurements, not faults.
What to do about it
- Fix one tool and one machine, then run it five times and take the median. Google states that the median of 5 runs is twice as stable as 1 run.
- Read
benchmarkIndexin both reports before you blame the page. - Record the Lighthouse version beside every score you keep.
- Set your pass threshold against a distribution, not one run. Google recommends the median, the 90th percentile, or min and max. If that threshold gates a build, the defaults matter — see how to run Lighthouse in CI.
What we do about it
Every Lightscore run pins the Lighthouse and Chromium version. The runtime line on the homepage names both, and each result repeats the version that produced it. All regions run on one machine class.
The API accepts up to 5 runs per region and returns the median run. The free one-URL tool runs once. That's a single sample, and the table above says what a single sample is worth.
We also hand back every run, not only the median one.
checks.lighthouse.per_run carries each run's four category scores next to its
benchmark_index. You can measure our noise yourself rather than trust this page
about it.
What we can't remove is your page. If your site runs an A/B test, or loads a tag on a timer, our numbers move too. No lab tool fixes that. A tool that promises identical scores every time hides the spread instead of measuring it.
Common questions
How many points of Lighthouse variance is normal?
On a fixed machine with pinned versions, most runs land within a few points of the median. Across 130 of our own single runs, 80 percent were within 2 points and 96 percent were within 5. The worst single run sat 14 points off. A busy page varies more than a light one.
Why is my local Lighthouse score higher than PageSpeed Insights?
Lighthouse grades a page partly on the speed of the machine that measured it, and the two machines are rarely the same. Different throttling settings and a different network path add to the gap. Open both report files and compare the benchmarkIndex value before you change any code.
Does running Lighthouse five times fix the problem?
It fixes noise, not bias. Google states that the median score of 5 runs is twice as stable as a single run. It does nothing about a systematic offset, such as a different Lighthouse version or a different CPU throttling setting.
Why do two tools on the same Lighthouse version still disagree?
Lighthouse reports what it measured on the hardware it ran on, under the throttling profile it was given. Two tools with the same version, different hardware and different throttling settings will report different numbers for the same page. The version is one input of several.