How to monitor Lighthouse scores over time
Pick one URL, one region, one device and one pinned Lighthouse version. Run that combination on a schedule. Store all five performance metrics beside the score, never the score on its own. Then alert when a metric leaves its own historical range and stays out.
Scheduling takes an afternoon. The alert is the part that goes wrong, because a Lighthouse score moves on its own, and a fixed threshold can't tell the difference.
Pin the setup, or you compare nothing
A trend line only means something if every point measured the same thing. Five inputs have to stay constant:
- The URL, query string included.
- The region the run started from.
- The device profile — mobile or desktop.
- The Lighthouse version.
- The Chromium version and the throttling profile.
Change one and the step in your chart belongs to your tooling. Scoring curves and metric weights change between Lighthouse releases, so an unpinned version rewrites your history without anyone touching the site.
Then store this much for every run:
- The timestamp, in UTC.
- All five performance metrics, in milliseconds — not just the score.
- The four category scores.
- The Lighthouse and Chromium versions.
environment.benchmarkIndexfrom the report JSON.- Your deploy or commit id, if you have one.
benchmarkIndex is the one people skip. It's Lighthouse's own rough measure of how
fast the host machine was, and it sits at the bottom of every report under
"CPU/Memory Power". Google's brackets put 1,500 to 2,000 at high-end desktop.
Without it you can't separate a slower page from a busier machine.
How much does the score move on its own?
We keep three example reports warm, so we have a fixed rig and a log of it. What
follows is wikipedia.org on mobile, measured from Washington, on Lighthouse
12.8.2, 41 times between 12 June and 13 August 2026. The median gap between runs
is about 32 hours.
The first 38 runs scored between 87 and 97, with a median of 92.5. Nothing in the setup changed. Across those runs Total Blocking TimeHow long the main thread was busy and couldn’t respond to taps or clicks. ranged from 16 ms to 398 ms, and Largest Contentful PaintTime until the largest thing in view (hero image, headline) has painted. from 2,128 ms to 2,541 ms.
That's a ten-point band on the score. Any alert threshold inside it fires on an ordinary day. If you want the longer treatment of where that spread comes from, we wrote it up separately in why Lighthouse scores differ between runs.
How do you tell a real regression from noise?
Look at the last three points: 78, 80, 79. Three questions, in this order.
Did it stay? One run under the band is a sample. Three consecutive runs under it are a state. This is the cheapest test and it removes most false alarms on its own.
Did the metrics move with it? Total Blocking Time went from a median of 235 ms to 688, 614 and 685 ms. Largest Contentful Paint went from a median of 2,295 ms to 2,667, 2,696 and 2,588 ms. Every one of those six readings sits outside the range the same page held for the previous 57 days. A score that drops while every metric stays put is a rounding artefact, not a regression.
Was the machine normal? The three runs reported a benchmarkIndex of 1,635,
1,719 and 1,750. The previous 38 runs spanned 1,589 to 2,066 on the same machine
class. So the host was ordinary, and the page got slower.
You can also check the arithmetic, which is the part that convinced me. Lighthouse gives Total Blocking Time a weight of 30 and scores it on a log-normal curve with control points at 200 ms and 600 ms. Moving the median from 235 ms to 685 ms costs 12.7 of those 30 points. The LCP shift costs another 1.7 of its 25. Predicted drop: about 14 points. Measured drop: 13.5.
When the arithmetic matches the observation that closely, you're looking at the page, not at the weather.
What to alert on instead of a score threshold
"Alert when performance drops below 90" is the rule most tutorials reach for. On our own two pages it tells you nothing, for two different reasons:
| Page | Runs | Score range | A "below 90" alert |
|---|---|---|---|
| wikipedia.org | 38, before 10 August | 87–97 | fires on an ordinary day |
| vercel.com | 40 | 30–36 | fires permanently, so it says nothing |
Vercel is the harder case. On 11 August its Total Blocking Time reached 9,072 ms, against a median of 2,521 ms. That's 3.6 times worse. The score went from 31 to 32. On that part of the curve a 6.5-second regression is worth 1.4 points. No score threshold anywhere could separate that day from an ordinary one.
So alert on the metrics, against the page's own history:
- Keep a rolling median and range per metric over your last 20 or 30 runs.
- Fire when a metric sits outside that range for three consecutive runs.
- Read the composite score. Don't page anyone on it.
Google's own guidance points the same way. Its notes on Lighthouse variability tell you to build thresholds from aggregates, not single results: the median, the 90th percentile, or min and max. The scoring documentation goes further. It suggests you treat site performance as a distribution of scores rather than one number.
Build the loop
The mechanism is a scheduled job, an audit API and somewhere to put the rows. Here it is against our free tool, which needs no key and no signup:
# 1. start a run
ID=$(curl -s -X POST https://lightscore.dev/report \
-d url=https://example.com/ -d regions=iad -d device=mobile -d format=json \
| jq -r .id)
# 2. poll until it finishes
while [ "$(curl -s https://lightscore.dev/report/example.com/$ID/status \
| jq -r .status)" != completed ]; do sleep 5; done
# 3. append one row: metrics, versions, machine
curl -s https://lightscore.dev/report/example.com/$ID.json | jq -c '{
at: .completed_at,
lighthouse: .lighthouse_version,
scores: .results[0].scores,
metrics: .results[0].metrics,
bench: .results[0].checks.lighthouse.per_run[0].benchmark_index
}' >> history.ndjson
Put that on a daily cron, add a second trigger on deploy, and you have the data half. Daily is also the useful ceiling here. Our free tool caches a result for 24 hours per URL, region set and device. An hourly cron just gets you yesterday's run back.
Three ways to get there, and the right one depends on what you already run:
| Approach | Pick it when | What it costs you |
|---|---|---|
| Cron against an audit API | you want the raw JSON in your own store | you build the chart and the alert |
| Lighthouse CI server | you already run lhci in your pipeline |
your runners vary, so the machine is a variable |
| A monitoring suite | you want dashboards and alerting today | a monthly subscription |
The middle row is worth spelling out. Lighthouse CI's server stores the runs your own CI produced. Those ran on whatever hardware the job landed on, which reintroduces the variable you pinned in the first section. It's a good store for running Lighthouse in CI and a poor one for a trend line.
What a lab monitor can't tell you
It can't measure Interaction to Next PaintHow long the page takes to visibly respond to a tap, click or key press.. A run is one scripted page load with nobody in front of it. INP needs a real person tapping something. Lighthouse gives it a weight of zero in the performance category for that reason.
It isn't your users. One region, one device profile and one throttling setting describe one visitor you invented. Field data from the Chrome UX Report describes the real ones. Run both, and expect them to disagree.
It can't see a page that changes underneath it. A/B tests, ad slots and third-party tags all move the numbers without a deploy. That shows up as a step in your chart and it isn't a regression you can fix.
what’s this
The small set of field metrics Google treats as a ranking input and reports in Search Console. We measure their lab equivalents here — real-user field data (CrUX) is a separate signal. TBT is the lab stand-in for INP, not a Core Web Vital itself.
One last piece of transparency, since this page is on our own site. Lightscore
doesn't schedule anything for you. There's no monitor to create, no trend chart and
no alerting — we're the measurement layer, and the loop above is genuinely the loop.
If you want the scheduling and the charts handed to you, DebugBear, Calibre and
SpeedCurve all sell that, and they're good at it. We give you the other half.
Every run carries a pinned runtime, the region you asked for, its own
benchmark_index, and the full Lighthouse JSON to keep. That's what you need to
build the history yourself.
Common questions
How often should I run Lighthouse to monitor a page?
Once a day is enough for most sites, and adding runs per deploy is better than raising the daily cadence. What matters more than frequency is that every run uses the same URL, region, device and Lighthouse version. An hourly schedule on a drifting setup tells you less than a daily one on a pinned setup.
How much does a Lighthouse score change on its own?
On our fixed rig, wikipedia.org returned scores between 87 and 97 across 38 runs with no change we caused. A quiet page moves far less. Our own homepage returned 99 or 100 on every run in the same window. Measure your own page before you pick an alert level.
Should I alert when the performance score drops below 90?
No. A fixed threshold sits inside the noise band of some pages and outside the reach of others. On one of our monitored pages that alert would fire on an ordinary day. On another, Total Blocking Time tripled and the score never came close to 90 in either direction. Alert on the metrics instead, against the page's own history.
Can Lighthouse monitoring measure Interaction to Next Paint?
No. A Lighthouse run is one scripted page load with no user in front of it, and INP needs a real interaction. Lighthouse gives INP a weight of zero in the performance category. For INP you need field data, such as the Chrome UX Report.