Lightscore

How to monitor Lighthouse scores over time

Pick one URL, one region, one device and one pinned Lighthouse version. Run that combination on a schedule. Store all five performance metrics beside the score, never the score on its own. Then alert when a metric leaves its own historical range and stays out.

Scheduling takes an afternoon. The alert is the part that goes wrong, because a Lighthouse score moves on its own, and a fixed threshold can't tell the difference.

Pin the setup, or you compare nothing

A trend line only means something if every point measured the same thing. Five inputs have to stay constant:

Change one and the step in your chart belongs to your tooling. Scoring curves and metric weights change between Lighthouse releases, so an unpinned version rewrites your history without anyone touching the site.

Then store this much for every run:

benchmarkIndex is the one people skip. It's Lighthouse's own rough measure of how fast the host machine was, and it sits at the bottom of every report under "CPU/Memory Power". Google's brackets put 1,500 to 2,000 at high-end desktop. Without it you can't separate a slower page from a busier machine.

How much does the score move on its own?

We keep three example reports warm, so we have a fixed rig and a log of it. What follows is wikipedia.org on mobile, measured from Washington, on Lighthouse 12.8.2, 41 times between 12 June and 13 August 2026. The median gap between runs is about 32 hours.

Performance score Total Blocking Time (ms) Days since 12 June 2026 0 20 40 60 80 100 0 100 200 300 400 500 600 700 0 10 20 30 40 50 60 70 Performance Total Blocking Time
One page, one region, one pinned Lighthouse build, 41 runs on a true 0–100 axis. The score holds its band for 57 days, then steps down exactly where Total Blocking Time steps up.

The first 38 runs scored between 87 and 97, with a median of 92.5. Nothing in the setup changed. Across those runs Total Blocking TimeHow long the main thread was busy and couldn’t respond to taps or clicks. ranged from 16 ms to 398 ms, and Largest Contentful PaintTime until the largest thing in view (hero image, headline) has painted. from 2,128 ms to 2,541 ms.

That's a ten-point band on the score. Any alert threshold inside it fires on an ordinary day. If you want the longer treatment of where that spread comes from, we wrote it up separately in why Lighthouse scores differ between runs.

How do you tell a real regression from noise?

Look at the last three points: 78, 80, 79. Three questions, in this order.

Did it stay? One run under the band is a sample. Three consecutive runs under it are a state. This is the cheapest test and it removes most false alarms on its own.

Did the metrics move with it? Total Blocking Time went from a median of 235 ms to 688, 614 and 685 ms. Largest Contentful Paint went from a median of 2,295 ms to 2,667, 2,696 and 2,588 ms. Every one of those six readings sits outside the range the same page held for the previous 57 days. A score that drops while every metric stays put is a rounding artefact, not a regression.

Was the machine normal? The three runs reported a benchmarkIndex of 1,635, 1,719 and 1,750. The previous 38 runs spanned 1,589 to 2,066 on the same machine class. So the host was ordinary, and the page got slower.

You can also check the arithmetic, which is the part that convinced me. Lighthouse gives Total Blocking Time a weight of 30 and scores it on a log-normal curve with control points at 200 ms and 600 ms. Moving the median from 235 ms to 685 ms costs 12.7 of those 30 points. The LCP shift costs another 1.7 of its 25. Predicted drop: about 14 points. Measured drop: 13.5.

When the arithmetic matches the observation that closely, you're looking at the page, not at the weather.

What to alert on instead of a score threshold

"Alert when performance drops below 90" is the rule most tutorials reach for. On our own two pages it tells you nothing, for two different reasons:

Page Runs Score range A "below 90" alert
wikipedia.org 38, before 10 August 87–97 fires on an ordinary day
vercel.com 40 30–36 fires permanently, so it says nothing

Vercel is the harder case. On 11 August its Total Blocking Time reached 9,072 ms, against a median of 2,521 ms. That's 3.6 times worse. The score went from 31 to 32. On that part of the curve a 6.5-second regression is worth 1.4 points. No score threshold anywhere could separate that day from an ordinary one.

So alert on the metrics, against the page's own history:

  1. Keep a rolling median and range per metric over your last 20 or 30 runs.
  2. Fire when a metric sits outside that range for three consecutive runs.
  3. Read the composite score. Don't page anyone on it.

Google's own guidance points the same way. Its notes on Lighthouse variability tell you to build thresholds from aggregates, not single results: the median, the 90th percentile, or min and max. The scoring documentation goes further. It suggests you treat site performance as a distribution of scores rather than one number.

Build the loop

The mechanism is a scheduled job, an audit API and somewhere to put the rows. Here it is against our free tool, which needs no key and no signup:

# 1. start a run
ID=$(curl -s -X POST https://lightscore.dev/report \
  -d url=https://example.com/ -d regions=iad -d device=mobile -d format=json \
  | jq -r .id)

# 2. poll until it finishes
while [ "$(curl -s https://lightscore.dev/report/example.com/$ID/status \
  | jq -r .status)" != completed ]; do sleep 5; done

# 3. append one row: metrics, versions, machine
curl -s https://lightscore.dev/report/example.com/$ID.json | jq -c '{
  at: .completed_at,
  lighthouse: .lighthouse_version,
  scores: .results[0].scores,
  metrics: .results[0].metrics,
  bench: .results[0].checks.lighthouse.per_run[0].benchmark_index
}' >> history.ndjson

Put that on a daily cron, add a second trigger on deploy, and you have the data half. Daily is also the useful ceiling here. Our free tool caches a result for 24 hours per URL, region set and device. An hourly cron just gets you yesterday's run back.

Three ways to get there, and the right one depends on what you already run:

Approach Pick it when What it costs you
Cron against an audit API you want the raw JSON in your own store you build the chart and the alert
Lighthouse CI server you already run lhci in your pipeline your runners vary, so the machine is a variable
A monitoring suite you want dashboards and alerting today a monthly subscription

The middle row is worth spelling out. Lighthouse CI's server stores the runs your own CI produced. Those ran on whatever hardware the job landed on, which reintroduces the variable you pinned in the first section. It's a good store for running Lighthouse in CI and a poor one for a trend line.

What a lab monitor can't tell you

It can't measure Interaction to Next PaintHow long the page takes to visibly respond to a tap, click or key press.. A run is one scripted page load with nobody in front of it. INP needs a real person tapping something. Lighthouse gives it a weight of zero in the performance category for that reason.

It isn't your users. One region, one device profile and one throttling setting describe one visitor you invented. Field data from the Chrome UX Report describes the real ones. Run both, and expect them to disagree.

It can't see a page that changes underneath it. A/B tests, ad slots and third-party tags all move the numbers without a deploy. That shows up as a step in your chart and it isn't a regression you can fix.

what’s this

The small set of field metrics Google treats as a ranking input and reports in Search Console. We measure their lab equivalents here — real-user field data (CrUX) is a separate signal. TBT is the lab stand-in for INP, not a Core Web Vital itself.

One last piece of transparency, since this page is on our own site. Lightscore doesn't schedule anything for you. There's no monitor to create, no trend chart and no alerting — we're the measurement layer, and the loop above is genuinely the loop. If you want the scheduling and the charts handed to you, DebugBear, Calibre and SpeedCurve all sell that, and they're good at it. We give you the other half. Every run carries a pinned runtime, the region you asked for, its own benchmark_index, and the full Lighthouse JSON to keep. That's what you need to build the history yourself.

Common questions