cargo perf --repeat N: report a median and a spread, not a point sample #295

Closed
opened 2026-09-10 00:47:39 +00:00 by viberfox-agent · 5 comments
Collaborator

Problem

cargo perf measures each scenario once and reports that single reading as the value. perf::Window (crates/cartopolis/src/systems/dev/perf.rs:297-347) has one f64/u64 per column, Report::windows is a flat Vec<Window> (perf.rs:366-370), and the harness flies each scenario exactly once — ShotProgress::index advances past every shot and never revisits one (crates/cartopolis/src/systems/dev/shot_harness.rs:2497-2508).

That is the shape of every result this suite has had to withdraw:

  • docs/notes/render-performance-survey.md:346-349 — "repeats of one config spanned 218–319 ms … The within-group spread exceeds every between-group difference." Three configurations were compared on medians the author had to compute by hand from three separate launches.
  • docs/notes/performance-suite.md:141-149 lists three withdrawn findings, each one sample differenced against another sample.
  • docs/notes/performance-suite.md:158 — "identical builds have been measured 25 % apart."

The module already knows this and already fixed it inside one window: GpuPass::frag_min/frag_max/frag_samples sample the fill counter every frame precisely because "a point sample of this is not reproducible and will mislead you" (perf.rs:283-293, sampled at perf.rs:573-581). Nothing does the same across windows, so every column that is not fragment_invocations is still a point sample, and the only way to get a spread today is to launch cargo perf N times and diff N JSON files by hand.

There is no --repeat anywhere in the client (grep repeat over perf.rs, shot_harness.rs, lib.rs returns only an unrelated comment at lib.rs:2688).

Nothing outside the crate reads perf.json — no test, no CI job, no tool (grep -rn "perf.json\|perf::Report" finds only docs and lib.rs:275). The report shape can change freely.

Approach

Two files, plus documentation.

1. crates/cartopolis/src/lib.rs — the flag

  • Add #[arg(long, default_value_t = 1, value_name = "N")] repeat: u32 beside the existing --perf-out / --perf-frames (lib.rs:113-120).
  • --repeat 0 is an error, not a clamp — same rule the manifest already applies to a typo'd --only name (perf.rs:158-162).
  • --repeat with a value other than 1 and no --perf is an error, following the precedent at lib.rs:299-301 where --only outside a --shots run is refused rather than ignored.
  • Pass it into systems::perf::build at lib.rs:268-277.

2. crates/cartopolis/src/systems/dev/perf.rs

Flying the repeats. build (perf.rs:444-525) already turns the selected scenarios into ShotHarness::shots and a parallel PerfRun { names, steady_frames }. Repeat all three lists repeat times round-robin (a b c a b c a b c), not blocked (a a a b b b). Reason: the drift this ticket exists to expose is the shared box, and blocked repeats put all three samples of one scenario adjacent in time, so a slow patch lands entirely inside one scenario's group and is attributed to the scene. Round-robin spreads it over all of them. Everything downstream already works per shot index: place_camera re-issues the teleport for whatever progress.index names (shot_harness.rs:2059-2072), and drive_perf_window reads progress.shot_index() to label the window (perf.rs:559, perf.rs:670-675). The end-of-run test perf.index + 1 >= run.names.len() (perf.rs:620) still fires on the last flight.

Add repeat: u32 to PerfRun (perf.rs:195-203) so finish can record it.

PNG filenames. ShotSpec::out is perf-{name}.png (perf.rs:473) and repeats would collide on one file. When repeat == 1 keep that name byte-for-byte; otherwise write perf-{name}-r{n}.png, 1-based. A picture per repeat is what explains a spread that turns out to be "a tile had not arrived that time" — the same reason perf.rs:470-472 gives for writing a PNG at all.

The report. Keep Window (perf.rs:297-347) exactly as it is — it is what close builds and what log_window prints per flight (perf.rs:718, perf.rs:723-769), so the raw per-run numbers stay in the log. Change what finish (perf.rs:772-806) serialises:

/// One numeric column over the repeats of one window.
pub struct Stat { pub runs: u32, pub min: f64, pub median: f64, pub max: f64 }

pub struct WindowSummary {
    pub scenario: String,
    pub phase: String,
    /// Flights aggregated here.
    pub runs: u32,
    /// How many of them passed `Window::quiet`.
    pub quiet_runs: u32,
    pub frames: Stat, pub wall_s: Stat,
    pub fps: Option<Stat>, pub frame_time_ms: Option<Stat>,
    pub process_cpu_percent: Option<Stat>, pub process_mem_percent: Option<Stat>,
    pub entity_count: Option<Stat>, pub mesh_slabs: Option<Stat>, pub mesh_slab_mb: Option<Stat>,
    pub upload_kb: Stat, pub worker_jobs: Stat, pub mesh_added_per_frame: Stat,
    pub resident_kb: Stat, pub resident_image_kb: Stat, pub resident_mesh_kb: Stat,
    pub draws: Stat, pub tris_total: Stat,
    pub gpu_passes: Vec<GpuPassSummary>,
}

pub struct GpuPassSummary {
    pub path: String,
    pub runs: u32,
    pub gpu_ms: Option<Stat>, pub cpu_ms: Option<Stat>,
    pub fragment_invocations: Option<Stat>,
    /// Min of the per-run mins, max of the per-run maxes, and the total frames sampled.
    pub frag_min: Option<f64>, pub frag_max: Option<f64>, pub frag_samples: u32,
}

pub struct Report { pub meta: ReportMeta, pub windows: Vec<WindowSummary> }

ReportMeta (perf.rs:350-364) gains pub repeat: u32.

The aggregation is one pure function — fn summarise(samples: &[Window]) -> Vec<WindowSummary> — taking PerfProgress::samples and grouping by (scenario, phase). Being pure and taking a slice is what makes it testable with no app (the whole of the rest of this module needs a World). Rules, all pinned by tests:

  • Group key is (scenario, phase); output order is first appearance, i.e. manifest order with load before steady for each scenario, so the JSON reads the same as a --repeat 1 report does today.
  • WindowSummary::runs is the number of flights in the group. Stat::runs is the number of those flights that carried a value for that column — an Option<f64> column absent in some runs aggregates over the present ones and says so, rather than counting a missing diagnostic as zero (the perf.rs:26-27 rule: "where they are missing the columns come back absent rather than zero"). A column absent in every run is None.
  • Median: sort ascending; odd n is the middle sample; even n is the mean of the two middle samples. The ordinary definition, chosen so a reader recomputing the median by hand from the perf window log lines gets the same number the JSON reports — a suite whose arithmetic disagrees with the reader's is another way to withdraw a result. The alternative (lower middle, so every reported number is one actually measured) is named here so nobody has to re-derive the choice; --repeat 3 never reaches the difference.
  • n == 1 gives min == median == max by construction, not by a special case.
  • gpu_passes merge by path. frag_min = min of the per-run frag_mins, frag_max = max of the per-run frag_maxes, frag_samples = sum. fragment_invocations, gpu_ms, cpu_ms become Stats. Sort descending by median fragment_invocations, matching the existing sort at perf.rs:425-429, with passes carrying no statistics last.

One summary log line per aggregated window, written from finish alongside the existing perf: wrote report line, in the shape log_window already uses (perf.rs:723-769) — scenario, phase, runs, quiet_runs, and median (min..max) for frame_time_ms and fps. Same reason that function gives: a shell that never opens the JSON still learns something.

3. Documentation

  • perf.rs module doc "Running it" (perf.rs:49-55) — add cargo perf --repeat 3 --only orbit.
  • docs/perf.toml header (lines 3-5) and the cargo perf comment block in .cargo/config.toml:83-88 — same line.
  • docs/notes/performance-suite.md — this is a standing fact, so it goes in the note (value 7, Say what you decided). Under "A point sample of the fill counter is not reproducible" (performance-suite.md:71-81), record that the same argument applies to every other column across launches, that --repeat N is the answer, the even-n median rule, and the caveat below.

The caveat that must be written down: a repeat's load window is not a cold one. Tiles are cached in platform::storage::Namespace::Cache on disk and in TileCache in memory (crates/cartopolis/src/systems/map/tile_loader.rs:143-146), so repeat 2 of a scenario re-streams from a warm process. load medians therefore describe a re-teleport, not a cold start, and repeat 1 is not comparable with repeats 2..N. (The disk half of this is already true today between launches; the in-memory half is new with --repeat.) The steady window is unaffected — it is what this flag is for.

Acceptance criteria

  • --repeat N exists on the CLI, defaults to 1, and is documented in its clap doc comment.
  • --repeat 0 fails at startup with a message naming the flag; --repeat other than 1 without --perf fails at startup, matching lib.rs:299-301.
  • cargo perf --repeat 3 --only orbit flies orbit three times in one launch and writes one report.
  • That report's windows has exactly two entries (orbit/load, orbit/steady), each with runs: 3, and each numeric column an object with runs, min, median, max.
  • meta.repeat records the repeat count.
  • cargo perf --only orbit (no flag) still produces two windows with runs: 1 and min == median == max on every column, and still writes perf-orbit.png under that exact name.
  • With repeat > 1, each flight writes its own perf-{name}-r{n}.png.
  • finish logs one summary line per aggregated window carrying runs, quiet_runs and the median plus range for frame_time_ms and fps; the existing per-flight perf window lines are unchanged.
  • Unit tests over synthetic Window values, all in perf.rs's existing #[cfg(test)] mod tests (perf.rs:819):
    • three samples of one (scenario, phase) → one summary, runs == 3, median is the middle value, min/max correct;
    • one sample → min == median == max, runs == 1;
    • four samples → median is the mean of the two middle values;
    • a column present in 2 of 3 runs → Stat::runs == 2 and the aggregate is over those two; a column None in all runs → the field is None;
    • two scenarios × two phases → four summaries in first-appearance order;
    • gpu_passes merge by path: frag_min is the min of the mins, frag_max the max of the maxes, frag_samples the sum, and a pass present in only some runs reports the smaller runs.
  • docs/notes/performance-suite.md, docs/perf.toml, .cargo/config.toml and the perf.rs module doc all mention --repeat, and the note carries the warm-load caveat.
  • The commit body names the direction value applied (value 1, Measured beats plausible — a median with a spread beside it is the difference between a measurement and a coincidence) and records the round-robin ordering decision and the even-n median rule.

Verification

Runs here, in this container (lavapipe; orbit is the one scenario cheap enough — docs/perf.toml:91-95 says a city view takes minutes to settle):

cargo fmt -p cartopolis -- --check
cargo test -p cartopolis perf
cargo perf --only orbit --repeat 3 --perf-out /tmp/perf-r3.json
cargo perf --only orbit            --perf-out /tmp/perf-r1.json

Then read /tmp/perf-r3.json: two windows, runs: 3, three distinct samples visible as min != max on frame_time_ms, and meta.repeat == 3. Read /tmp/perf-r1.json: runs: 1 and min == median == max throughout. Cross-check the medians against the three perf window log lines the run printed.

The timing columns in either report are properties of a CPU rasteriser on a shared box and are not evidence of anything about the client (docs/notes/performance-suite.md:155-159) — they are being read here only to confirm the aggregation ran, not to conclude anything.

Only a workstation or the Steam Deck runner can check that the spread the flag now reports is small enough to make a between-config difference readable — the question docs/notes/render-performance-survey.md:344-352 could not answer. That is a measurement, so it is not this ticket's; it is what the next A/B uses the flag for. Nothing in this pass was compiled: this was a read-only refinement.

Out of scope

  • Raw per-run windows in the JSON. windows carries the aggregate only. Every individual flight is already printed by log_window (perf.rs:718), so nothing is lost; adding a samples array is a follow-up if a reader ever wants n > 5 raw.
  • A standard deviation, a confidence interval, or an outlier rule. Min/median/max is what the ticket asks for and what a reader can act on.
  • Gating on the spread — no --expect-style assertion over a perf report. There is no CI job running cargo perf at all today.
  • A repeat key in docs/perf.toml. CLI only; the manifest's deny_unknown_fields (perf.rs:83) will reject one, which is the correct answer.
  • Making load cold between repeats (cache eviction, a fresh process per repeat). Documented as a caveat, not solved.
  • Running the suite on Android, still unbuilt (docs/notes/performance-suite.md:161-163).
  • Adding scenarios to docs/perf.toml. An entry there is a claim the viewpoint has been measured (docs/perf.toml:91-95).

Open questions

None.


Branch: feat/295-perf-repeat-median-spread

Original request

From the maintainer's focus for this lane:

Improving performance on the steamdeck. keep it concrete and simple. use or extend harness tooling to find useful improvements.

cargo perf --repeat N: report a median and a spread, not a point sample

Lands in crates/cartopolis/src/systems/dev/perf.rs and the --perf flag block in lib.rs. Every withdrawn result in docs/notes/performance-suite.md was one sample differenced against one other sample, and docs/notes/render-performance-survey.md records a within-group spread wider than every between-group difference it was trying to measure. --repeat N re-flies the selected scenarios inside one launch and each Window in the JSON carries runs, min, median and max for every numeric column instead of a single value. Judged by: a --repeat 3 report whose windows carry three samples and a median, plus a unit test over synthetic windows pinning the aggregation, including that --repeat 1 still reports min == median == max. Needs no Steam Deck to build or test.

This is item 2 of 10 of the lane's queue. It is ordered by a person in the
admin console, so build this one rather than the one you would have picked — and keep it
the size it is. A ticket that turns out to be three tickets is one ticket: finish this
part and say on it what the other two are.

If the item is simply wrong — already done, impossible, or a bad idea against the values
— say so here and close it. That is a result, not a failure.

Decide it yourself. This ticket is not being watched, so a question asked here is a
ticket that stops. The seven values at the top of docs/direction.md exist to settle
exactly that kind of ambiguity: pick the reading they support, say in the commit body
which one you applied, and build. Only a decision needing something nobody can derive
from the repository — a credential, a licence somebody must accept, a choice about what
the project is for — is a reason to stop.

Filed by the autopilot.

🤖 Refined by the viberfox issue agent. Reply with @agent refine and what is wrong to have this rewritten.

## Problem `cargo perf` measures each scenario **once** and reports that single reading as the value. `perf::Window` (`crates/cartopolis/src/systems/dev/perf.rs:297-347`) has one `f64`/`u64` per column, `Report::windows` is a flat `Vec<Window>` (`perf.rs:366-370`), and the harness flies each scenario exactly once — `ShotProgress::index` advances past every shot and never revisits one (`crates/cartopolis/src/systems/dev/shot_harness.rs:2497-2508`). That is the shape of every result this suite has had to withdraw: - `docs/notes/render-performance-survey.md:346-349` — "repeats of one config spanned 218–319 ms … **The within-group spread exceeds every between-group difference.**" Three configurations were compared on medians the author had to compute by hand from three separate launches. - `docs/notes/performance-suite.md:141-149` lists three withdrawn findings, each one sample differenced against another sample. - `docs/notes/performance-suite.md:158` — "identical builds have been measured 25 % apart." The module already knows this and already fixed it **inside** one window: `GpuPass::frag_min`/`frag_max`/`frag_samples` sample the fill counter every frame precisely because "a point sample of this is not reproducible and will mislead you" (`perf.rs:283-293`, sampled at `perf.rs:573-581`). Nothing does the same **across** windows, so every column that is not `fragment_invocations` is still a point sample, and the only way to get a spread today is to launch `cargo perf` N times and diff N JSON files by hand. There is no `--repeat` anywhere in the client (`grep repeat` over `perf.rs`, `shot_harness.rs`, `lib.rs` returns only an unrelated comment at `lib.rs:2688`). Nothing outside the crate reads `perf.json` — no test, no CI job, no tool (`grep -rn "perf.json\|perf::Report"` finds only docs and `lib.rs:275`). The report shape can change freely. ## Approach Two files, plus documentation. ### 1. `crates/cartopolis/src/lib.rs` — the flag - Add `#[arg(long, default_value_t = 1, value_name = "N")] repeat: u32` beside the existing `--perf-out` / `--perf-frames` (`lib.rs:113-120`). - `--repeat 0` is an error, not a clamp — same rule the manifest already applies to a typo'd `--only` name (`perf.rs:158-162`). - `--repeat` with a value other than 1 and no `--perf` is an error, following the precedent at `lib.rs:299-301` where `--only` outside a `--shots` run is refused rather than ignored. - Pass it into `systems::perf::build` at `lib.rs:268-277`. ### 2. `crates/cartopolis/src/systems/dev/perf.rs` **Flying the repeats.** `build` (`perf.rs:444-525`) already turns the selected scenarios into `ShotHarness::shots` and a parallel `PerfRun { names, steady_frames }`. Repeat all three lists `repeat` times **round-robin** (`a b c a b c a b c`), not blocked (`a a a b b b`). Reason: the drift this ticket exists to expose is the shared box, and blocked repeats put all three samples of one scenario adjacent in time, so a slow patch lands entirely inside one scenario's group and is attributed to the scene. Round-robin spreads it over all of them. Everything downstream already works per shot index: `place_camera` re-issues the teleport for whatever `progress.index` names (`shot_harness.rs:2059-2072`), and `drive_perf_window` reads `progress.shot_index()` to label the window (`perf.rs:559`, `perf.rs:670-675`). The end-of-run test `perf.index + 1 >= run.names.len()` (`perf.rs:620`) still fires on the last flight. Add `repeat: u32` to `PerfRun` (`perf.rs:195-203`) so `finish` can record it. **PNG filenames.** `ShotSpec::out` is `perf-{name}.png` (`perf.rs:473`) and repeats would collide on one file. When `repeat == 1` keep that name byte-for-byte; otherwise write `perf-{name}-r{n}.png`, 1-based. A picture per repeat is what explains a spread that turns out to be "a tile had not arrived that time" — the same reason `perf.rs:470-472` gives for writing a PNG at all. **The report.** Keep `Window` (`perf.rs:297-347`) exactly as it is — it is what `close` builds and what `log_window` prints per flight (`perf.rs:718`, `perf.rs:723-769`), so the raw per-run numbers stay in the log. Change what `finish` (`perf.rs:772-806`) serialises: ```rust /// One numeric column over the repeats of one window. pub struct Stat { pub runs: u32, pub min: f64, pub median: f64, pub max: f64 } pub struct WindowSummary { pub scenario: String, pub phase: String, /// Flights aggregated here. pub runs: u32, /// How many of them passed `Window::quiet`. pub quiet_runs: u32, pub frames: Stat, pub wall_s: Stat, pub fps: Option<Stat>, pub frame_time_ms: Option<Stat>, pub process_cpu_percent: Option<Stat>, pub process_mem_percent: Option<Stat>, pub entity_count: Option<Stat>, pub mesh_slabs: Option<Stat>, pub mesh_slab_mb: Option<Stat>, pub upload_kb: Stat, pub worker_jobs: Stat, pub mesh_added_per_frame: Stat, pub resident_kb: Stat, pub resident_image_kb: Stat, pub resident_mesh_kb: Stat, pub draws: Stat, pub tris_total: Stat, pub gpu_passes: Vec<GpuPassSummary>, } pub struct GpuPassSummary { pub path: String, pub runs: u32, pub gpu_ms: Option<Stat>, pub cpu_ms: Option<Stat>, pub fragment_invocations: Option<Stat>, /// Min of the per-run mins, max of the per-run maxes, and the total frames sampled. pub frag_min: Option<f64>, pub frag_max: Option<f64>, pub frag_samples: u32, } pub struct Report { pub meta: ReportMeta, pub windows: Vec<WindowSummary> } ``` `ReportMeta` (`perf.rs:350-364`) gains `pub repeat: u32`. The aggregation is one pure function — `fn summarise(samples: &[Window]) -> Vec<WindowSummary>` — taking `PerfProgress::samples` and grouping by `(scenario, phase)`. Being pure and taking a slice is what makes it testable with no app (the whole of the rest of this module needs a `World`). Rules, all pinned by tests: - Group key is `(scenario, phase)`; output order is **first appearance**, i.e. manifest order with `load` before `steady` for each scenario, so the JSON reads the same as a `--repeat 1` report does today. - `WindowSummary::runs` is the number of flights in the group. `Stat::runs` is the number of those flights that carried a value for **that column** — an `Option<f64>` column absent in some runs aggregates over the present ones and says so, rather than counting a missing diagnostic as zero (the `perf.rs:26-27` rule: "where they are missing the columns come back absent rather than zero"). A column absent in every run is `None`. - Median: sort ascending; odd n is the middle sample; **even n is the mean of the two middle samples**. The ordinary definition, chosen so a reader recomputing the median by hand from the `perf window` log lines gets the same number the JSON reports — a suite whose arithmetic disagrees with the reader's is another way to withdraw a result. The alternative (lower middle, so every reported number is one actually measured) is named here so nobody has to re-derive the choice; `--repeat 3` never reaches the difference. - `n == 1` gives `min == median == max` by construction, not by a special case. - `gpu_passes` merge by `path`. `frag_min` = min of the per-run `frag_min`s, `frag_max` = max of the per-run `frag_max`es, `frag_samples` = sum. `fragment_invocations`, `gpu_ms`, `cpu_ms` become `Stat`s. Sort descending by median `fragment_invocations`, matching the existing sort at `perf.rs:425-429`, with passes carrying no statistics last. **One summary log line per aggregated window**, written from `finish` alongside the existing `perf: wrote report` line, in the shape `log_window` already uses (`perf.rs:723-769`) — scenario, phase, runs, quiet_runs, and `median (min..max)` for `frame_time_ms` and `fps`. Same reason that function gives: a shell that never opens the JSON still learns something. ### 3. Documentation - `perf.rs` module doc "Running it" (`perf.rs:49-55`) — add `cargo perf --repeat 3 --only orbit`. - `docs/perf.toml` header (lines 3-5) and the `cargo perf` comment block in `.cargo/config.toml:83-88` — same line. - `docs/notes/performance-suite.md` — this is a standing fact, so it goes in the note (value 7, *Say what you decided*). Under "A point sample of the fill counter is not reproducible" (`performance-suite.md:71-81`), record that the same argument applies to every other column across launches, that `--repeat N` is the answer, the even-n median rule, and the caveat below. **The caveat that must be written down:** a repeat's `load` window is not a cold one. Tiles are cached in `platform::storage::Namespace::Cache` on disk and in `TileCache` in memory (`crates/cartopolis/src/systems/map/tile_loader.rs:143-146`), so repeat 2 of a scenario re-streams from a warm process. `load` medians therefore describe a re-teleport, not a cold start, and repeat 1 is not comparable with repeats 2..N. (The disk half of this is already true today between launches; the in-memory half is new with `--repeat`.) The `steady` window is unaffected — it is what this flag is for. ## Acceptance criteria - [ ] `--repeat N` exists on the CLI, defaults to 1, and is documented in its clap doc comment. - [ ] `--repeat 0` fails at startup with a message naming the flag; `--repeat` other than 1 without `--perf` fails at startup, matching `lib.rs:299-301`. - [ ] `cargo perf --repeat 3 --only orbit` flies `orbit` three times in one launch and writes one report. - [ ] That report's `windows` has exactly two entries (`orbit`/`load`, `orbit`/`steady`), each with `runs: 3`, and each numeric column an object with `runs`, `min`, `median`, `max`. - [ ] `meta.repeat` records the repeat count. - [ ] `cargo perf --only orbit` (no flag) still produces two windows with `runs: 1` and `min == median == max` on every column, and still writes `perf-orbit.png` under that exact name. - [ ] With `repeat > 1`, each flight writes its own `perf-{name}-r{n}.png`. - [ ] `finish` logs one summary line per aggregated window carrying runs, quiet_runs and the median plus range for `frame_time_ms` and `fps`; the existing per-flight `perf window` lines are unchanged. - [ ] Unit tests over synthetic `Window` values, all in `perf.rs`'s existing `#[cfg(test)] mod tests` (`perf.rs:819`): - [ ] three samples of one `(scenario, phase)` → one summary, `runs == 3`, median is the middle value, min/max correct; - [ ] one sample → `min == median == max`, `runs == 1`; - [ ] four samples → median is the mean of the two middle values; - [ ] a column present in 2 of 3 runs → `Stat::runs == 2` and the aggregate is over those two; a column `None` in all runs → the field is `None`; - [ ] two scenarios × two phases → four summaries in first-appearance order; - [ ] `gpu_passes` merge by path: `frag_min` is the min of the mins, `frag_max` the max of the maxes, `frag_samples` the sum, and a pass present in only some runs reports the smaller `runs`. - [ ] `docs/notes/performance-suite.md`, `docs/perf.toml`, `.cargo/config.toml` and the `perf.rs` module doc all mention `--repeat`, and the note carries the warm-`load` caveat. - [ ] The commit body names the direction value applied (value 1, *Measured beats plausible* — a median with a spread beside it is the difference between a measurement and a coincidence) and records the round-robin ordering decision and the even-n median rule. ## Verification Runs here, in this container (lavapipe; `orbit` is the one scenario cheap enough — `docs/perf.toml:91-95` says a city view takes minutes to settle): ```bash cargo fmt -p cartopolis -- --check cargo test -p cartopolis perf cargo perf --only orbit --repeat 3 --perf-out /tmp/perf-r3.json cargo perf --only orbit --perf-out /tmp/perf-r1.json ``` Then read `/tmp/perf-r3.json`: two windows, `runs: 3`, three distinct samples visible as `min != max` on `frame_time_ms`, and `meta.repeat == 3`. Read `/tmp/perf-r1.json`: `runs: 1` and `min == median == max` throughout. Cross-check the medians against the three `perf window` log lines the run printed. The timing columns in either report are properties of a CPU rasteriser on a shared box and are not evidence of anything about the client (`docs/notes/performance-suite.md:155-159`) — they are being read here only to confirm the aggregation ran, not to conclude anything. **Only a workstation or the Steam Deck runner can check** that the spread the flag now reports is small enough to make a between-config difference readable — the question `docs/notes/render-performance-survey.md:344-352` could not answer. That is a measurement, so it is not this ticket's; it is what the next A/B uses the flag for. Nothing in this pass was compiled: this was a read-only refinement. ## Out of scope - **Raw per-run windows in the JSON.** `windows` carries the aggregate only. Every individual flight is already printed by `log_window` (`perf.rs:718`), so nothing is lost; adding a `samples` array is a follow-up if a reader ever wants n > 5 raw. - **A standard deviation, a confidence interval, or an outlier rule.** Min/median/max is what the ticket asks for and what a reader can act on. - **Gating on the spread** — no `--expect`-style assertion over a perf report. There is no CI job running `cargo perf` at all today. - **A `repeat` key in `docs/perf.toml`.** CLI only; the manifest's `deny_unknown_fields` (`perf.rs:83`) will reject one, which is the correct answer. - **Making `load` cold between repeats** (cache eviction, a fresh process per repeat). Documented as a caveat, not solved. - **Running the suite on Android**, still unbuilt (`docs/notes/performance-suite.md:161-163`). - **Adding scenarios to `docs/perf.toml`.** An entry there is a claim the viewpoint has been measured (`docs/perf.toml:91-95`). ## Open questions None. --- Branch: `feat/295-perf-repeat-median-spread` <details><summary>Original request</summary> From the maintainer's focus for this lane: > Improving performance on the steamdeck. keep it concrete and simple. use or extend harness tooling to find useful improvements. **cargo perf --repeat N: report a median and a spread, not a point sample** Lands in crates/cartopolis/src/systems/dev/perf.rs and the --perf flag block in lib.rs. Every withdrawn result in docs/notes/performance-suite.md was one sample differenced against one other sample, and docs/notes/render-performance-survey.md records a within-group spread wider than every between-group difference it was trying to measure. --repeat N re-flies the selected scenarios inside one launch and each Window in the JSON carries runs, min, median and max for every numeric column instead of a single value. Judged by: a --repeat 3 report whose windows carry three samples and a median, plus a unit test over synthetic windows pinning the aggregation, including that --repeat 1 still reports min == median == max. Needs no Steam Deck to build or test. This is item 2 of 10 of the lane's queue. It is ordered by a person in the admin console, so build this one rather than the one you would have picked — and keep it the size it is. A ticket that turns out to be three tickets is one ticket: finish this part and say on it what the other two are. If the item is simply wrong — already done, impossible, or a bad idea against the values — say so here and close it. That is a result, not a failure. **Decide it yourself.** This ticket is not being watched, so a question asked here is a ticket that stops. The seven values at the top of `docs/direction.md` exist to settle exactly that kind of ambiguity: pick the reading they support, say in the commit body which one you applied, and build. Only a decision needing something nobody can derive from the repository — a credential, a licence somebody must accept, a choice about what the project is *for* — is a reason to stop. <sub>Filed by the autopilot. <!-- autopilot:focus-cargo-perf-repeat-n-report-a-median-and --></sub> </details> <sub>🤖 Refined by the viberfox issue agent. Reply with **@agent refine** and what is wrong to have this rewritten.</sub>
Author
Collaborator

🤖 Refining this before anything gets built — reading the code, no branch, no commits. If it turns out I'd have to guess at the requirements I'll stop and ask instead.

🤖 Refining this before anything gets built — reading the code, no branch, no commits. If it turns out I'd have to guess at the requirements I'll stop and ask instead.
Author
Collaborator

🤖 Refined. The description above is now the spec; I'll pick it up for implementation on the next poll.

If the direction is wrong, say so now — @agent refine with a correction rewrites it, and agent:skip stops me.


🤖 viberfox issue agent on the VPS — 22 turns · $2.54 · session 53840b6a-ee46-4d5a-862a-5c34f525cb33. Reply with @agent to continue this session.

🤖 **Refined.** The description above is now the spec; I'll pick it up for implementation on the next poll. If the direction is wrong, say so now — **@agent refine** with a correction rewrites it, and `agent:skip` stops me. --- <sub>🤖 viberfox issue agent on the VPS — 22 turns · $2.54 · session `53840b6a-ee46-4d5a-862a-5c34f525cb33`. Reply with **@agent** to continue this session.</sub>
Author
Collaborator

🤖 Picked this up — working on branch feat/295-perf-repeat-median-spread in a fresh worktree on the VPS. I'll report back here when I'm done.

Label an issue agent:skip to keep me off it, or driven to keep it in a session you are driving.

🤖 Picked this up — working on branch `feat/295-perf-repeat-median-spread` in a fresh worktree on the VPS. I'll report back here when I'm done. <sub>Label an issue `agent:skip` to keep me off it, or `driven` to keep it in a session you are driving.</sub>
Author
Collaborator

The session was killed after hitting the 45 minute limit.

No commits were made.


🤖 viberfox issue agent on the VPS — . Reply with @agent to continue this session.

The session was killed after hitting the 45 minute limit. _No commits were made._ --- <sub>🤖 viberfox issue agent on the VPS — . Reply with **@agent** to continue this session.</sub>
Owner

Built in the nightly run: PR #306 (feat/295-perf-repeat), verify-branch green on e486d02, and checked end to end with cargo perf --only orbit --repeat 3. Not merged.

Built in the nightly run: PR #306 (`feat/295-perf-repeat`), verify-branch green on `e486d02`, and checked end to end with `cargo perf --only orbit --repeat 3`. Not merged.
Sign in to join this conversation.
No milestone
No project
No assignees
2 participants
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
jeroen/cartopolis#295
No description provided.