How to run Lighthouse in CI
Install @lhci/cli, write a lighthouserc.js, and run lhci autorun as a CI step.
That takes about ten minutes, and every guide on this query covers it.
The gate is the part that costs you a week. Two defaults decide what it really checks: how many times Lighthouse runs, and which of those runs meets your threshold. The shipped answers are one run, and the best one. This page sets Lighthouse CI up, then changes both.
The short version
Lighthouse CI does three things in order. It collects N Lighthouse runs, asserts them against your thresholds, and uploads the reports somewhere you can read them.
Put the config at the root of your repo:
// lighthouserc.js
module.exports = {
ci: {
collect: {
url: ['https://staging.example.com/'],
numberOfRuns: 5,
},
assert: {
preset: 'lighthouse:recommended',
aggregationMethod: 'median',
},
upload: {
target: 'temporary-public-storage',
},
},
};
Then call it from a job:
# .github/workflows/lighthouse.yml
name: lighthouse
on: pull_request
jobs:
audit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@@v7
- uses: treosh/lighthouse-ci-action@@v12
with:
configPath: ./lighthouserc.js
For a static build, swap url for staticDistDir: './dist' and Lighthouse CI
serves the directory itself. For an app that needs a server, use
startServerCommand: 'npm start'.
If GitHub Actions is your CI, the action has its own defaults and its own traps —
how to run Lighthouse in a GitHub Action covers
the runner you get and what it does to the numbers. On GitLab, the image most people
reach for holds two different Lighthouse versions —
how to run Lighthouse in GitLab CI has the working
.gitlab-ci.yml and the version trap.
Your gate checks one run, and the best one
The Lighthouse CI CLI runs each URL three times by default. The GitHub Action does not inherit that number. Its source reads:
runs: core.getInput('runs') ? parseInt(core.getInput('runs'), 10) : numberOfRuns || 1
So the action takes the runs input first, then numberOfRuns from your config
file, and falls back to 1. Set neither and your gate is a single sample.
The second default is the one I'd call a genuine trap. Lighthouse CI aggregates
several runs into one number before it compares against your threshold. The setting
is aggregationMethod. It takes median, optimistic, pessimistic or
median-run, and it defaults to optimistic.
Optimistic means most likely to pass. For a minScore assertion it takes the
highest score across the runs. For a maxNumericValue assertion it takes the
lowest number. Your build is graded on its luckiest run.
The two defaults stack in a way that hides the problem. With one run there's no
median to take, so nothing looks wrong. Raise numberOfRuns to fix the noise and
you silently switch on best-of-N grading instead. Set aggregationMethod: 'median'
explicitly. It is one line and it changes what the gate means.
Why does my Lighthouse CI job fail on a branch I didn't touch?
The runner. Lighthouse grades a page partly on the speed of the machine that measured it, and that speed is not a property of your code.
Google's own notes on Lighthouse variability set a floor of two dedicated cores and 2 GB of RAM, and recommend four cores and 4–8 GB. They tell you to avoid function-as-a-service platforms and burstable or shared-core instances. The troubleshooting guide is blunter: free CI providers "tend to have very underpowered CPUs".
A ubuntu-latest GitHub runner gives you 4 vCPU on a public repository and 2 vCPU
on a private one. A private repo therefore sits exactly on Google's minimum, not its
recommendation.
Default throttling doesn't rescue you. Lighthouse simulates a slow network rather than applying one. Google's mitigation table rates that as full mitigation for network variability, and only partial mitigation against client hardware and resource contention. The network is modelled. The CPU is real.
Here is what that costs, from our own fleet. We measured wikipedia.org on mobile
from Frankfurt and Washington, on one pinned Lighthouse build, seconds apart:
| Region | Performance | Total Blocking TimeHow long the main thread was busy and couldn’t respond to taps or clicks. | LCPTime until the largest thing in view (hero image, headline) has painted. |
|---|---|---|---|
| Frankfurt | 90 | 299 ms | 2,481 ms |
| Washington | 79 | 685 ms | 2,588 ms |
Eleven points apart, same page, same build. LCP moved by 107 ms. Total Blocking Time more than doubled. Total Blocking Time counts main-thread work, so it's the number a slower machine hits hardest. You can check both figures in the public report for that run.
A CI runner is one more machine in that spread. Your threshold has to survive it.
What does lighthouse:recommended actually check?
lighthouse:recommended is the preset almost every tutorial hands you, and it's a
reasonable starting point. It is also more cautious than its reputation.
Nine audits sit in the preset's source under a comment that reads // Flaky audits (warn). They are downgraded from error to warning, so they cannot fail your build:
first-contentful-paint,largest-contentful-paint,speed-indexcumulative-layout-shift,interactive,max-potential-fidbootup-time,mainthread-work-breakdown,first-meaningful-paint
That's most of the performance section, including two {term:core_web_vitals|Core Web
Vitals}. total-blocking-time isn't on the list because the base preset sets it to
off altogether.
So a green build under the recommended preset can't fail on LCPTime until the largest thing in view (hero image, headline) has painted., CLSHow much the page jumps around as it loads. Lower is better; 0 is rock-steady. or Total Blocking Time. You still get the warnings in the log. Nothing stops the merge. Google shipped it that way on purpose, and the comment says why.
Set a gate that holds
Look at what the preset does still fail on, and a pattern shows up. Those checks
measure a fact rather than a conclusion. unminified-javascript either found
unminified files or it didn't. uses-text-compression either saw compression or it
didn't. Neither answer moves when the runner is busy.
Lighthouse CI's troubleshooting guide names this directly. Assert on facts such as JavaScript request counts and sizes, not on conclusions such as Time to Interactive.
Four changes, in the order I'd make them:
- Set
numberOfRuns: 5. Google states the median of 5 runs is twice as stable as 1. - Set
aggregationMethod: 'median'on theassertblock. - Keep the score assertions at
warn. Read them, don't merge on them. - Add
maxNumericValueassertions on byte counts and request counts, aterror.
assert: {
preset: 'lighthouse:recommended',
aggregationMethod: 'median',
assertions: {
'total-byte-weight': ['error', {maxNumericValue: 500000}],
'unused-javascript': ['error', {maxNumericValue: 100000}],
'categories:performance': ['warn', {minScore: 0.8}],
},
},
A budget in bytes fails when someone adds a dependency. It doesn't fail because the runner was busy. That is the whole difference between a gate people trust and a gate people learn to re-run.
what’s this
The small set of field metrics Google treats as a ranking input and reports in Search Console. We measure their lab equivalents here — real-user field data (CrUX) is a separate signal. TBT is the lab stand-in for INP, not a Core Web Vital itself.
What a CI run can't tell you
Three things, and no configuration fixes any of them.
Your build isn't your site. Lighthouse CI works best against localhost or a
preview deploy, which is exactly why it fits a pre-merge gate. That page has no CDN
in front of it, no production cache, and often no real third-party tags. Your
visitors get none of those conditions.
One machine isn't the world. The table above is two of ours. A runner in
us-east-1 tells you what a page costs from us-east-1.
Version skew is silent. @lhci/cli 0.15.1 bundles Lighthouse 12.6.1. If another
tool in your pipeline runs a different Lighthouse version, the two aren't comparable,
because scoring curves and metric weights change between releases. Record the version
next to every score you keep.
The honest split we'd suggest: keep Lighthouse CI as the pre-merge check on your own build, where byte budgets catch real regressions early. Then check the deployed URL separately, after the deploy, from the places your users are — that half is monitoring Lighthouse scores over time, and it needs a different threshold rule than a CI gate does.
That second half is what Lightscore does. You POST a URL and pick regions. The
scores come back with the full Lighthouse JSON and HTML, on a pinned Lighthouse and
Chromium build. There's a worked curl example in
how to run an audit from CI or a script. The API takes a runs parameter up
to 5 and returns the median run, with every individual run's scores alongside it.
What we can't do is audit your localhost or a preview URL behind auth. A worker in
Frankfurt has no route to your build agent. For that stage, Lighthouse CI is the
right tool and we aren't a replacement for it.
Common questions
How do I run Lighthouse in CI?
Install the @lhci/cli package, add a lighthouserc.js file with your URLs and assertions, then run lhci autorun as a CI step. The command collects several Lighthouse runs, asserts them against your thresholds, and uploads the reports. On GitHub Actions you can use the treosh/lighthouse-ci-action wrapper instead of calling the CLI yourself.
Why does my Lighthouse CI job fail on a branch that did not change?
Almost always the runner, not the code. Lighthouse grades a page partly on the speed of the machine that measured it, and CI runners vary. Google recommends at least two dedicated cores and warns that free CI environments are volatile. Raise numberOfRuns and assert against the median instead of a single result.
How many Lighthouse runs should CI do?
Five is a reasonable default. Google states that the median score of 5 runs is twice as stable as 1 run. The Lighthouse CI CLI defaults to 3. The GitHub Action falls back to 1 when neither its runs input nor numberOfRuns in your config file sets a value.
Should CI fail the build on a low performance score?
Not on the composite score alone. It moves with the runner hardware, so it produces false failures. Assert on facts the run measures directly instead — total JavaScript bytes, unminified files, missing compression. Keep the score as a warning you read rather than a gate that blocks a merge.