Skip to content

Comment benchmark results on pull requests - #1529

Open
jprendes wants to merge 12 commits into
hyperlight-dev:mainfrom
jprendes:benchmark-prs
Open

jprendes wants to merge 12 commits into
hyperlight-dev:mainfrom
jprendes:benchmark-prs

Conversation

@jprendes

@jprendes jprendes commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

Every pull request gets a benchmark comment: one table per configuration, compared against the commit the branch left main at, with both commits named.

cargo ci

hyperlight-ci is a small binary, aliased to cargo ci, holding the benchmark plumbing that used to be shell in the Justfile and the workflows.

  • cargo ci bench builds and runs the suite.
  • cargo ci bench-report renders the markdown.

A report reads results from a directory or straight from GitHub, so the comment can be reproduced locally:

--candidate run:<id> | pr:<number> | commit:<sha> | base-of:<number> | release:<tag>

--baseline takes the same forms and defaults to the branch point. Runs are paired to their baseline by os, cpu vendor and hypervisor. Downloads cache under ci-runs.

Each run uploads its criterion results and keeps them 90 days. A pull request is compared against the run at its branch point, so that run has to outlive the branch.

Running them without wrecking the numbers

Benchmarks run in parallel across physical cores, which is what makes them affordable on every pull request. Each benchmark is pinned to its own core, and the physical core owning CPU 0 is left out because it carries the timer and the interrupt work. A sandbox is held resident for the whole run, so nothing crosses the boundary between zero and one live VM. KVM toggles a static key there, and each toggle patches kernel text and IPIs every core, landing on whichever benchmark happens to trigger it.

bench_report.toml says which benchmarks are steady enough to report and how far a result moves before it counts as a change. A pool answers with whichever machine is free, so two runs of the same commit often land on different processors: Linux_mshv3_amd spreads 1.6% within one EPYC generation and 8.9% across the two generations its pool holds. The thresholds are deliberately generous to absorb that. Measured over eleven runs of one commit, the defaults of 1.1x and 0.9x produce an alert in every report, while the committed values report clean across all 55 pairings. They can tighten as more data comes in, or if a pool is pinned to one generation.

@jprendes
jprendes force-pushed the benchmark-prs branch 2 times, most recently from 8ad9393 to 0e181c4 Compare June 12, 2026 10:43
@jprendes jprendes added the kind/enhancement For PRs adding features, improving functionality, docs, tests, etc. label Jun 12, 2026
@jprendes
jprendes force-pushed the benchmark-prs branch 4 times, most recently from bc499ea to a277ac1 Compare June 17, 2026 14:48
@jprendes
jprendes force-pushed the benchmark-prs branch 3 times, most recently from e660953 to 9ce9ed0 Compare June 18, 2026 11:22
@hyperlight-gh-bot

This comment has been minimized.

@hyperlight-gh-bot

This comment has been minimized.

@hyperlight-dev hyperlight-dev deleted a comment from hyperlight-gh-bot Bot Jun 18, 2026
@hyperlight-gh-bot

This comment has been minimized.

@hyperlight-gh-bot

This comment has been minimized.

@hyperlight-gh-bot

This comment has been minimized.

@jprendes
jprendes marked this pull request as ready for review June 19, 2026 15:54
Copilot AI review requested due to automatic review settings June 19, 2026 15:54
@jprendes
jprendes force-pushed the benchmark-prs branch 5 times, most recently from 67e75dd to 6fa38df Compare September 25, 2026 19:00
@hyperlight-gh-bot

This comment has been minimized.

@hyperlight-gh-bot

This comment has been minimized.

@hyperlight-gh-bot

This comment has been minimized.

@hyperlight-gh-bot

This comment has been minimized.

Introduce a new internal tooling crate (hyperlight-ci) that provides:

- bench subcommand: Runs criterion benchmarks in parallel via
  criterion-swarm. Features include:
  - Configurable parallelism (-j N, defaults to all P-cores)
  - Configurable output modes (spinner, stream, summary)
  - Support for pre-built binaries (--binary) to skip rebuilds
  - Trailing args forwarded to criterion (filter, --exact, etc.)

- bench-report subcommand: Generates markdown comparison tables from
  criterion's target/criterion/ JSON output via criterion-markdown.
  Features include:
  - Benchmark discovery via criterion-swarm
  - Optional allowlist filtering via --binary or trailing args
  - Output to stdout

This replaces ad-hoc benchmark scripting with a unified tool suitable
for both local development and CI report generation.

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
- Add cargo alias (`cargo ci`) for convenient hyperlight-ci invocation
- Update dep_benchmarks workflow to use `cargo ci bench` and generate
  a markdown report via `cargo ci bench-report`, posting results as
  a PR comment per hypervisor/cpu matrix entry
- Add benchmarks job to ValidatePullRequest workflow with hypervisor
  and cpu matrix, gated behind docs-only and build-guests checks
- Grant pull-requests: write permission for PR comment posting
- Simplify Justfile bench recipes to delegate to `cargo ci bench`
- Update benchmarking docs to reflect the new workflow

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
`bench_report.toml` lists regular expressions matched against criterion benchmark ids. A benchmark is selected when it matches `allowlist` and no `denylist` entry, and an empty `allowlist` keeps everything the denylist does not exclude. Omitting both lists disables filtering. `cargo ci bench-report --config-file` reports the selection, and `cargo ci bench --config-file` runs it. CI keeps running every benchmark and filters at report time.

Both subcommands apply the selection through `CriterionSwarm::retain`, added in criterion-swarm 0.2.1, so running and reporting cannot drift apart.

An allowlist pattern that matches no benchmark fails, so a rename surfaces instead of dropping out of the comment silently. A denylist pattern matching nothing is accepted, because a benchmark may be absent on some platforms. Unknown keys are rejected so a stale key cannot silently disable filtering.

The initial lists keep 47 of 116 benchmarks: those whose median drifted by at most 5% across five back-to-back runs on an idle machine, less the snapshot cold start and restore families.

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
A vCPU created without an in-kernel LAPIC bumps the kernel's
`kvm_has_noapic_vcpu` static key, and teardown drops it again. Each
transition through zero rewrites kernel text and IPIs every core. Benchmarks
that create and drop sandboxes leave no resident VM between iterations, so
they cross that boundary constantly: a full suite run issues 176k broadcast
IPIs, against 560 with one sandbox held.

`cargo ci bench` now keeps one sandbox alive for the duration of a run, so
that cost lands on neither the benchmark that triggers it nor its neighbours.
Excluding benchmarks that ran on CPU 0, which has its own much larger effect,
median drift for the sandbox group falls from 6.5% to 1.7%. Pass
`--no-ballast` to measure the cold path instead.

The helper is an example rather than a dependency of this crate, keeping
hyperlight-host out of the CI tool's build. It exits on stdin EOF, so it
cannot outlive the run even when this process is killed without unwinding.

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
Criterion stores an id per result directory but nothing about the run as a
whole, and results accumulate: a directory keeps benchmarks that no longer
exist, indistinguishable from the ones just measured. CI compounds this by
unpacking a baseline into the same directory before running, so its uploaded
artifact holds 155 result directories for a suite of 87.

`cargo ci bench` writes `benchmarks.json` next to the results, listing the
benchmarks the run covers along with a timestamp and the host it ran on. The
list is the post-filter set, so it describes what was measured rather than
what was discovered.

`bench-report` takes the ids from there when the results carry one, so a
criterion directory from elsewhere, a CI artifact say, reads without a
toolchain and without the checkout matching. Naming them by listing the
binaries builds them first and describes the current checkout instead, which
is the same thing only while reporting a local run of the current tree.
Rendering the downloaded results of a full suite takes 11ms rather than a
build.

Explicit `--binary` or trailing bench args still list the binaries, since both
name benchmarks the manifest cannot filter, as do results from before the
manifest existed.

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
Comparing what CI measured meant downloading six artifacts by hand, unpacking
each somewhere, and running the report once per configuration. The results are
richer than the comment CI posts, holding every benchmark rather than the
reported subset and the samples behind each estimate, so reaching for them is
worth making cheap.

`--candidate` and `--baseline` say where each side comes from: a criterion
directory, `run:<ID>`, or `pr:<NUMBER>`, which takes the most recent run that
still has its artifacts, since the newest is often a label check or one whose
benchmarks have not finished. A whole run renders as one section per
hypervisor and cpu vendor, the way the pull request comment reads. Runs land
under `target/ci-runs` and are reused, artifacts being immutable. A run is
about 400MB unpacked, and the first report of one waits on the download.

criterion-markdown 0.2.1 computes changes when it renders, so the report needs
both datasets present. A run overwrites the baseline it compares against, so
CI keeps the previous results in their own criterion root. Criterion keeps the
last run of a directory in `new` and the one before it in `base`, so a
directory on its own reports the last run against the previous one, and
another directory is compared through its own last run. Results with no
baseline to compare against are reported on their own.

CI covers every hypervisor and cpu vendor, and comparing results measured on
different machines says nothing. Each side is paired with the one that ran on
the same kind of machine, read from the artifact name or from the operating
system, cpu vendor and hypervisor a run records. What fits nothing, or
several, is reported without a comparison.

The ids come from the manifest, so they describe what the run measured rather
than this checkout, which need not even be the same commit.

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
The default branch is benchmarked daily rather than per commit, so the run
that measured a given commit rarely exists. `commit:<SHA>` takes the closest
run that carries no changes the commit never had, and `base-of:<NUMBER>` the
one where a pull request branched. Nothing within a pull request's own
results says what they mean, so that is what they are measured against. A
cancelled run leaves some configurations unmeasured, so only whole runs
serve as a baseline.

Sections are named the way the workflow that measured them is, so a report
of a run reads like the comment CI posts.

Artifacts hold nothing every run is bound to leave behind, so a download
that finished says so itself. Reaching for a file the run might not have
written re-downloaded results already on disk, onto the ones already there.

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
Each benchmark job rendered its own section against a baseline it
downloaded itself, which only the pull request comment ever read. The daily
and release runs rendered sections nothing consumes, and a pull request was
measured against the latest release rather than the branch it targets.

The job that posts the comment now reports the whole run against where the
pull request branched, so benchmarking a configuration is only that. What
the results are worth on their own outlives the comparison, so a baseline
out of reach costs the changes, not the report.

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
Workflow artifacts were swept away after five days, so comparing against
anything older meant unpacking a release tarball by hand into a directory
the report could be pointed at. They now last as long as GitHub keeps them,
and `release:<TAG>` reads what a release carries, which outlives artifacts
altogether. Searching as far back as results are kept lets a commit that old
still be found.

Benchmark ids change over time, so a release far enough back has little left
to compare against. What matches is reported and the rest stands alone.

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
A comment that only holds numbers leaves the reader to work out what was
measured and what it was held against. The report names both commits, and
`--reproduce` ends it with the command that asks for it again. What was
asked for moves, since the last run of a pull request is whichever ran most
recently and where it branched changes when it is rebased, so the command
names the runs that answered instead.

GitHub answers run listings out of an index that takes a moment to warm,
leaving the most recent runs out of the first replies. Taking one at its
word picked a baseline months older than the one asked for, silently,
because an old run is still an ancestor of the commit. Listings are now
asked for until two agree on the newest run.

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
The lists were measured on a quiet machine, which is not the machine that
reports them, and then from four runs whose spread was taken as the range
between the extremes. Whole runs land slow: two of eleven sat 10% to 20%
above the rest across the entire suite, and a range reads that as every
benchmark being unstable. The middle half of the runs says what a benchmark
usually does.

Listed here when no configuration spreads more than 10%. That is a fraction
of what any one configuration could carry. `Linux_kvm_amd` holds 87 of 95
benchmarks within 5%, `guest_calls` among them at 0.3%, while
`Windows_hyperv-ws2025_amd` holds 15 and stays there on uniform hardware.
`Linux_mshv3_amd` is as quiet as kvm on one processor and four times worse
across two, because its pool answers with both EPYC generations.

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
What the report calls a regression is a judgement about the machine it ran
on, not about the benchmark. Two runs of one commit raise an alert in every
report at the default ratios, and in none of them once a change has to reach
1.5x either way.

`improvement`, `strong_improvement` and `regression` set where those lines
fall, in the config file beside the benchmarks they apply to, or on the
command line for a one-off reading. The renderer rejects ratios that cross
or invert. `summary_limit` and `reproduce` settle beside them, being what a
repository decides once rather than per report.

`repo` reads `remote:<name>` as whichever repository a git remote points at,
so a fork reads its own runs instead of the one written into the tool.

The config file is read once for a report rather than once for each
configuration it covers.

Signed-off-by: Jorge Prendes <jorge.prendes@gmail.com>
@hyperlight-gh-bot

Copy link
Copy Markdown

Benchmark Results

Measured commit: bc678968f3e8
Baseline commit: 5131e8332f88

kvm / amd (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 775.63 ns (➖ 1.09x faster)
vec_bytes 578.56 ns (➖ 1.01x slower)
376.42 µs (➖ 1.01x slower)

payload_allocation

slot_pool_segmented
262144 523.25 ns (➖ 1.02x faster)
65536 142.88 ns (➖ 1.01x slower)

sandboxes

create_initialized_and_drop
medium 78.25 ms (➖ 1.03x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
7.77 ns (➖ 1.00x slower) 7.90 ns (➖ 1.01x slower) 7.76 ns (➖ 1.08x faster)

snapshot_files

load_snapshot_unverified
small 91.98 µs (➖ 1.01x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.21 µs (➖ 1.13x faster) 7.34 µs (➖ 1.02x slower)
65536 2.08 µs (➖ 1.08x slower) 2.06 µs (➖ 1.07x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 6.26 µs (➖ 1.10x faster) 6.26 µs (➖ 1.04x faster)
8192 1.05 µs (➖ 1.00x faster) 1.07 µs (➖ 1.00x slower)
262144 27.21 µs (➖ 1.05x faster)
kvm / intel (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 711.95 ns (➖ 1.32x faster)
vec_bytes 509.59 ns (➖ 1.17x faster)
671.64 µs (➖ 1.20x faster)

payload_allocation

slot_pool_segmented
262144 506.86 ns (➖ 1.11x faster)
65536 136.61 ns (➖ 1.11x faster)

sandboxes

create_initialized_and_drop
medium 79.09 ms (➖ 1.24x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
6.92 ns (➖ 1.11x faster) 6.92 ns (➖ 1.20x faster) 6.94 ns (➖ 1.10x faster)

snapshot_files

load_snapshot_unverified
small 44.61 µs (➖ 1.10x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.57 µs (➖ 1.11x faster) 7.61 µs (➖ 1.10x faster)
65536 2.14 µs (➖ 1.09x faster) 2.15 µs (➖ 1.09x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.13 µs (➖ 1.23x faster) 7.13 µs (➖ 1.24x faster)
8192 726.51 ns (➖ 1.22x faster) 726.52 ns (➖ 1.19x faster)
262144 29.65 µs (➖ 1.19x faster)
mshv3 / amd (Linux) (❌ 1.83x)

Top regressions

  • sandboxes/create_initialized_and_drop/medium — ❌ 1.83x slower

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 963.17 ns (➖ 1.04x faster)
vec_bytes 713.53 ns (➖ 1.06x slower)
343.96 µs (➖ 1.00x slower)

payload_allocation

slot_pool_segmented
262144 713.64 ns (➖ 1.00x faster)
65536 194.58 ns (➖ 1.01x slower)

sandboxes

create_initialized_and_drop
medium 55.68 ms (❌ 1.83x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.77 ns (➖ 1.00x slower) 9.71 ns (➖ 1.00x faster) 10.67 ns (➖ 1.04x slower)

snapshot_files

load_snapshot_unverified
small 84.24 µs (➖ 1.10x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 9.61 µs (➖ 1.01x slower) 9.84 µs (➖ 1.08x slower)
65536 2.45 µs (➖ 1.02x slower) 2.24 µs (➖ 1.06x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.20 µs (➖ 1.09x faster) 8.22 µs (➖ 1.02x faster)
8192 1.29 µs (➖ 1.03x slower) 1.33 µs (➖ 1.08x slower)
262144 36.14 µs (➖ 1.07x faster)
mshv3 / intel (Linux) (❌ 2.48x)

Top regressions

  • function_call_codec/decode_vec_bytes_copy — ❌ 2.48x slower
  • sandboxes/create_initialized_and_drop/medium — ❌ 2.04x slower

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 984.69 ns (➖ 1.02x slower)
vec_bytes 681.18 ns (➖ 1.11x slower)
1.78 ms (❌ 2.48x slower)

payload_allocation

slot_pool_segmented
262144 639.38 ns (➖ 1.08x slower)
65536 169.65 ns (➖ 1.09x slower)

sandboxes

create_initialized_and_drop
medium 79.52 ms (❌ 2.04x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
8.84 ns (➖ 1.10x slower) 8.84 ns (➖ 1.09x slower) 8.87 ns (➖ 1.10x slower)

snapshot_files

load_snapshot_unverified
small 51.35 µs (➖ 1.16x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.94 µs (➖ 1.09x faster) 8.12 µs (➖ 1.07x faster)
65536 2.32 µs (➖ 1.06x faster) 2.31 µs (➖ 1.06x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.72 µs (➖ 1.16x faster) 7.80 µs (➖ 1.14x faster)
8192 925.92 ns (➖ 1.06x slower) 903.18 ns (➖ 1.02x faster)
262144 40.26 µs (➖ 1.10x slower)
hyperv-ws2025 / amd (Windows) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.52 µs (➖ 1.20x slower)
vec_bytes 783.14 ns (➖ 1.01x faster)
2.13 ms (➖ 1.17x slower)

payload_allocation

slot_pool_segmented
262144 799.77 ns (➖ 1.00x slower)
65536 237.26 ns (➖ 1.03x slower)

sandboxes

create_initialized_and_drop
medium 89.23 ms (➖ 1.38x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.94 ns (➖ 1.01x faster) 11.21 ns (➖ 1.12x slower) 10.14 ns (➖ 1.05x slower)

snapshot_files

load_snapshot_unverified
small 732.32 µs (➖ 1.09x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 9.00 µs (➖ 1.08x faster) 9.97 µs (➖ 1.04x slower)
65536 2.39 µs (➖ 1.01x faster) 2.40 µs (➖ 1.01x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 9.08 µs (➖ 1.01x slower) 8.94 µs (➖ 1.00x faster)
8192 1.30 µs (➖ 1.00x slower) 1.33 µs (➖ 1.03x slower)
262144 40.35 µs (➖ 1.01x slower)
hyperv-ws2025 / intel (Windows) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.19 µs (➖ 1.02x slower)
vec_bytes 840.05 ns (➖ 1.06x slower)
3.61 ms (➖ 1.22x slower)

payload_allocation

slot_pool_segmented
262144 744.81 ns (➖ 1.05x faster)
65536 212.04 ns (➖ 1.10x faster)

sandboxes

create_initialized_and_drop
medium 123.33 ms (➖ 1.36x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.18 ns (➖ 1.00x faster) 10.78 ns (➖ 1.00x slower) 10.19 ns (➖ 1.01x slower)

snapshot_files

load_snapshot_unverified
small 654.78 µs (➖ 1.35x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.88 µs (➖ 1.01x slower) 10.04 µs (➖ 1.29x slower)
65536 2.33 µs (➖ 1.01x slower) 2.31 µs (➖ 1.00x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.58 µs (➖ 1.01x faster) 8.44 µs (➖ 1.01x slower)
8192 1.15 µs (➖ 1.06x slower) 1.13 µs (➖ 1.02x faster)
262144 55.33 µs (➖ 1.42x slower)

Reported by cargo ci bench-report --candidate run:36475261745 --baseline run:36361996946 --config-file bench_report.toml.

@ludfjig

ludfjig commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Are benchmarks comment posted on all prs? Do we want an opt-in? Or maybe the point is to not have opt-in so we can't miss regression. Maybe opt-out for thigns like docs-only prs? Just some questnios :)

@jprendes

Copy link
Copy Markdown
Contributor Author

If the PR doesn't create a pr-comment artifact, it won't publish anything.
We skip that step for doc only PRs here

@jprendes jprendes changed the title Benchmark prs Comment benchmark results on pull requests Sep 28, 2026
@jprendes jprendes added area/infrastructure Concerns infrastructure rather than core functionality ready-for-review PR is ready for (re-)review labels Sep 28, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/infrastructure Concerns infrastructure rather than core functionality kind/enhancement For PRs adding features, improving functionality, docs, tests, etc. ready-for-review PR is ready for (re-)review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants