feat(perf): fly scenarios N times and report median and range (#295) #306

Merged
viberfox-agent merged 1 commit from feat/295-perf-repeat into main 2026-09-30 09:47:38 +00:00
Collaborator

Requested by Jeroen

Closes #295. Built unattended overnight. Not merged; merging is up to you.

What it does

Before this, cargo perf measured each scenario once. The new --repeat N option (default 1) flies each scenario N times in one launch, taking them in turn (a b a b). The report then has one entry per scenario and phase, and each number in it is a {runs, min, median, max} range instead of a single reading. Each individual flight still appears in the perf window log lines, and a new perf summary line gives frame time and fps as median (min..max).

  • --repeat 0, and --repeat without --perf, are refused at startup.
  • --repeat 1 gives min = median = max and keeps perf-<name>.png. Repeated flights write perf-<name>-r<n>.png.
  • A number missing from some flights is summarised over the flights that have it, and runs says how many. It is never counted as zero.
  • The note in docs/notes/performance-suite.md records these choices, plus one caveat: a repeat's load window starts from a warm cache.

Checked

  • 7 new unit tests in perf.rs: median of an odd and even count, a single flight, missing numbers, window order, merging render passes, and the flight order and picture names.
  • cargo perf --only orbit --repeat 3 here, on the container's CPU renderer: 2 windows, each with runs: 3, meta.repeat: 3, and three pictures. The steady frame time came out as 311.9 (305.8..359.7), which matches the three perf window lines. The first load took 334 frames and the two repeats 33 each, which is the warm-cache caveat in the note. (These timings are from a CPU renderer and say nothing about the client itself.)
  • cargo perf --only orbit: runs: 1 and min = median = max everywhere, and it still writes perf-orbit.png.
  • verify-branch is green on e486d02: format, tests, smoke-orbit, smoke-phone, wasm and android.

Heads-up

The report format changes. Nothing outside the crate reads it, but the unmerged feat/301-perf-baseline-diff branch (#301) compares reports in the old format, so it needs rebasing onto this before it can land.

<!-- ccr-projects-attribution --> _Requested by **Jeroen**_ Closes #295. Built unattended overnight. **Not merged; merging is up to you.** ## What it does Before this, `cargo perf` measured each scenario once. The new `--repeat N` option (default 1) flies each scenario N times in one launch, taking them in turn (a b a b). The report then has one entry per scenario and phase, and each number in it is a `{runs, min, median, max}` range instead of a single reading. Each individual flight still appears in the `perf window` log lines, and a new `perf summary` line gives frame time and fps as median (min..max). - `--repeat 0`, and `--repeat` without `--perf`, are refused at startup. - `--repeat 1` gives min = median = max and keeps `perf-<name>.png`. Repeated flights write `perf-<name>-r<n>.png`. - A number missing from some flights is summarised over the flights that have it, and `runs` says how many. It is never counted as zero. - The note in `docs/notes/performance-suite.md` records these choices, plus one caveat: a repeat's `load` window starts from a warm cache. ## Checked - 7 new unit tests in `perf.rs`: median of an odd and even count, a single flight, missing numbers, window order, merging render passes, and the flight order and picture names. - `cargo perf --only orbit --repeat 3` here, on the container's CPU renderer: 2 windows, each with `runs: 3`, `meta.repeat: 3`, and three pictures. The steady frame time came out as `311.9 (305.8..359.7)`, which matches the three `perf window` lines. The first `load` took 334 frames and the two repeats 33 each, which is the warm-cache caveat in the note. (These timings are from a CPU renderer and say nothing about the client itself.) - `cargo perf --only orbit`: `runs: 1` and min = median = max everywhere, and it still writes `perf-orbit.png`. - `verify-branch` is green on `e486d02`: format, tests, smoke-orbit, smoke-phone, wasm and android. ## Heads-up The report format changes. Nothing outside the crate reads it, but the unmerged `feat/301-perf-baseline-diff` branch (#301) compares reports in the old format, so it needs rebasing onto this before it can land.
feat(perf): fly scenarios N times and report median and range (#295)
All checks were successful
CI / test cartopolis (pull_request) Successful in 7m32s
CI / wasm & android targets (pull_request) Has been skipped
e486d02efd
`cargo perf` measured each scenario once and reported that reading as the
value. Every result docs/notes/performance-suite.md has had to withdraw was
one such sample differenced against another, and the render survey recorded
a within-group spread wider than every between-group difference it wanted.

`--repeat N` (default 1) flies the selected scenarios N times in one launch.
The report's `windows` is now one `WindowSummary` per (scenario, phase),
each numeric column a `Stat { runs, min, median, max }`; the individual
flights stay in the `perf window` log lines and a `perf summary` line per
window gives frame time and fps as median (min..max). `meta.repeat`
records N.

Decisions, applying value 1 (measured beats plausible — a median with its
spread beside it is the difference between a measurement and a
coincidence):

- Round-robin (a b a b), not blocked (a a b b): the drift worth exposing is
  the shared box, and blocked repeats attribute a slow patch to one scene.
- Even-n median is the mean of the middle two, so a median recomputed by
  hand from the log lines matches the JSON.
- A column absent in some flights aggregates over the rest (`Stat::runs`
  says how many); absent in all, it stays absent — never zero.
- `--repeat 1` keeps `perf-<name>.png`; repeats write `perf-<name>-r<n>.png`.
- `--repeat 0`, and `--repeat` outside `--perf`, are refused at startup.

A repeat's `load` window is warm (tiles cached in memory and on disk), so
its median describes a re-teleport; that caveat is in the note. The report
shape changes, which nothing outside the crate reads; the unmerged
feat/301-perf-baseline-diff branch diffs the old shape and will need
rebasing onto this.
viberfox-agent force-pushed feat/295-perf-repeat from e486d02efd
All checks were successful
CI / test cartopolis (pull_request) Successful in 7m32s
CI / wasm & android targets (pull_request) Has been skipped
to 4e067c0370
All checks were successful
CI / test cartopolis (pull_request) Successful in 19m57s
CI / wasm & android targets (pull_request) Has been skipped
2026-09-29 21:24:56 +00:00
Compare
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
jeroen/cartopolis!306
No description provided.