Lightscore

How to run Lighthouse in CI

Install @lhci/cli, write a lighthouserc.js, and run lhci autorun as a CI step. That takes about ten minutes, and every guide on this query covers it.

The gate is the part that costs you a week. Two defaults decide what it really checks: how many times Lighthouse runs, and which of those runs meets your threshold. The shipped answers are one run, and the best one. This page sets Lighthouse CI up, then changes both.

The short version

Lighthouse CI does three things in order. It collects N Lighthouse runs, asserts them against your thresholds, and uploads the reports somewhere you can read them.

Put the config at the root of your repo:

// lighthouserc.js
module.exports = {
  ci: {
    collect: {
      url: ['https://staging.example.com/'],
      numberOfRuns: 5,
    },
    assert: {
      preset: 'lighthouse:recommended',
      aggregationMethod: 'median',
    },
    upload: {
      target: 'temporary-public-storage',
    },
  },
};

Then call it from a job:

# .github/workflows/lighthouse.yml
name: lighthouse
on: pull_request
jobs:
  audit:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@@v7
      - uses: treosh/lighthouse-ci-action@@v12
        with:
          configPath: ./lighthouserc.js

For a static build, swap url for staticDistDir: './dist' and Lighthouse CI serves the directory itself. For an app that needs a server, use startServerCommand: 'npm start'.

If GitHub Actions is your CI, the action has its own defaults and its own traps — how to run Lighthouse in a GitHub Action covers the runner you get and what it does to the numbers. On GitLab, the image most people reach for holds two different Lighthouse versions — how to run Lighthouse in GitLab CI has the working .gitlab-ci.yml and the version trap.

Your gate checks one run, and the best one

The Lighthouse CI CLI runs each URL three times by default. The GitHub Action does not inherit that number. Its source reads:

runs: core.getInput('runs') ? parseInt(core.getInput('runs'), 10) : numberOfRuns || 1

So the action takes the runs input first, then numberOfRuns from your config file, and falls back to 1. Set neither and your gate is a single sample.

The second default is the one I'd call a genuine trap. Lighthouse CI aggregates several runs into one number before it compares against your threshold. The setting is aggregationMethod. It takes median, optimistic, pessimistic or median-run, and it defaults to optimistic.

Optimistic means most likely to pass. For a minScore assertion it takes the highest score across the runs. For a maxNumericValue assertion it takes the lowest number. Your build is graded on its luckiest run.

optimistic
(default)

median

lhci collect
numberOfRuns: 3

Scores
0.78 · 0.85 · 0.91

aggregationMethod
minScore: 0.9

Compares 0.91
passes

Compares 0.85
fails

Three runs, one threshold. The default aggregation compares the best score, not the middle one.

The two defaults stack in a way that hides the problem. With one run there's no median to take, so nothing looks wrong. Raise numberOfRuns to fix the noise and you silently switch on best-of-N grading instead. Set aggregationMethod: 'median' explicitly. It is one line and it changes what the gate means.

Why does my Lighthouse CI job fail on a branch I didn't touch?

The runner. Lighthouse grades a page partly on the speed of the machine that measured it, and that speed is not a property of your code.

Google's own notes on Lighthouse variability set a floor of two dedicated cores and 2 GB of RAM, and recommend four cores and 4–8 GB. They tell you to avoid function-as-a-service platforms and burstable or shared-core instances. The troubleshooting guide is blunter: free CI providers "tend to have very underpowered CPUs".

A ubuntu-latest GitHub runner gives you 4 vCPU on a public repository and 2 vCPU on a private one. A private repo therefore sits exactly on Google's minimum, not its recommendation.

Default throttling doesn't rescue you. Lighthouse simulates a slow network rather than applying one. Google's mitigation table rates that as full mitigation for network variability, and only partial mitigation against client hardware and resource contention. The network is modelled. The CPU is real.

Here is what that costs, from our own fleet. We measured wikipedia.org on mobile from Frankfurt and Washington, on one pinned Lighthouse build, seconds apart:

Region Performance Total Blocking TimeHow long the main thread was busy and couldn’t respond to taps or clicks. LCPTime until the largest thing in view (hero image, headline) has painted.
Frankfurt 90 299 ms 2,481 ms
Washington 79 685 ms 2,588 ms

Eleven points apart, same page, same build. LCP moved by 107 ms. Total Blocking Time more than doubled. Total Blocking Time counts main-thread work, so it's the number a slower machine hits hardest. You can check both figures in the public report for that run.

A CI runner is one more machine in that spread. Your threshold has to survive it.

What does lighthouse:recommended actually check?

lighthouse:recommended is the preset almost every tutorial hands you, and it's a reasonable starting point. It is also more cautious than its reputation.

Nine audits sit in the preset's source under a comment that reads // Flaky audits (warn). They are downgraded from error to warning, so they cannot fail your build:

That's most of the performance section, including two {term:core_web_vitals|Core Web Vitals}. total-blocking-time isn't on the list because the base preset sets it to off altogether.

So a green build under the recommended preset can't fail on LCPTime until the largest thing in view (hero image, headline) has painted., CLSHow much the page jumps around as it loads. Lower is better; 0 is rock-steady. or Total Blocking Time. You still get the warnings in the log. Nothing stops the merge. Google shipped it that way on purpose, and the comment says why.

Set a gate that holds

Look at what the preset does still fail on, and a pattern shows up. Those checks measure a fact rather than a conclusion. unminified-javascript either found unminified files or it didn't. uses-text-compression either saw compression or it didn't. Neither answer moves when the runner is busy.

Lighthouse CI's troubleshooting guide names this directly. Assert on facts such as JavaScript request counts and sizes, not on conclusions such as Time to Interactive.

Four changes, in the order I'd make them:

  1. Set numberOfRuns: 5. Google states the median of 5 runs is twice as stable as 1.
  2. Set aggregationMethod: 'median' on the assert block.
  3. Keep the score assertions at warn. Read them, don't merge on them.
  4. Add maxNumericValue assertions on byte counts and request counts, at error.
assert: {
  preset: 'lighthouse:recommended',
  aggregationMethod: 'median',
  assertions: {
    'total-byte-weight': ['error', {maxNumericValue: 500000}],
    'unused-javascript': ['error', {maxNumericValue: 100000}],
    'categories:performance': ['warn', {minScore: 0.8}],
  },
},

A budget in bytes fails when someone adds a dependency. It doesn't fail because the runner was busy. That is the whole difference between a gate people trust and a gate people learn to re-run.

what’s this

The small set of field metrics Google treats as a ranking input and reports in Search Console. We measure their lab equivalents here — real-user field data (CrUX) is a separate signal. TBT is the lab stand-in for INP, not a Core Web Vital itself.

What a CI run can't tell you

Three things, and no configuration fixes any of them.

Your build isn't your site. Lighthouse CI works best against localhost or a preview deploy, which is exactly why it fits a pre-merge gate. That page has no CDN in front of it, no production cache, and often no real third-party tags. Your visitors get none of those conditions.

One machine isn't the world. The table above is two of ours. A runner in us-east-1 tells you what a page costs from us-east-1.

Version skew is silent. @lhci/cli 0.15.1 bundles Lighthouse 12.6.1. If another tool in your pipeline runs a different Lighthouse version, the two aren't comparable, because scoring curves and metric weights change between releases. Record the version next to every score you keep.

The honest split we'd suggest: keep Lighthouse CI as the pre-merge check on your own build, where byte budgets catch real regressions early. Then check the deployed URL separately, after the deploy, from the places your users are — that half is monitoring Lighthouse scores over time, and it needs a different threshold rule than a CI gate does.

That second half is what Lightscore does. You POST a URL and pick regions. The scores come back with the full Lighthouse JSON and HTML, on a pinned Lighthouse and Chromium build. There's a worked curl example in how to run an audit from CI or a script. The API takes a runs parameter up to 5 and returns the median run, with every individual run's scores alongside it.

What we can't do is audit your localhost or a preview URL behind auth. A worker in Frankfurt has no route to your build agent. For that stage, Lighthouse CI is the right tool and we aren't a replacement for it.

Common questions