main is red on 407c1e5a #237

Closed
opened 2026-08-30 10:55:44 +00:00 by viberfox-agent · 6 comments
Collaborator

Problem

main is red at 407c1e5a because the Headless flicker smoke (street) step (.forgejo/workflows/ci.yml:459) was killed by its own timeout-minutes: 15 (.forgejo/workflows/ci.yml:461). Failing job: run 682, job 1259.

Nothing is wrong with the tests — every suite in that job reports test result: ok, 0 failed. The step started its cargo run at 09:49:33 and the runner logged ⚙️ [runner]: context deadline exceeded at 10:04:18, 14 m 45 s later. It had saved capture 1 of the flicker pair and was 6 minutes into capture 2.

This is not a flake, and it is not the runner. The step has been running at 60–95 % of its own cap since it landed (4cebab4, 2026‑08‑29). Every main run of it so far:

job commit capture 1 settle (virtual s) mean frame ms step wall clock
1234 c815861d 30.1 1968 8 m 53 s
1242 20b036a1 45.6 2340 13 m 26 s
1246 ab430c5a 30.0 1984 9 m 04 s
1250 29405f29 45.7 2669 13 m 36 s
1254 bdd05d07 45.9 2557 13 m 31 s
1259 407c1e5a 45.8 2551 killed at 14 m 45 s

Where the time actually goes

The second capture of a --flicker pair re-streams the entire city, for a 5 cm camera move.

--flicker turns each viewpoint into two shots, the twin nudged FLICKER_NUDGE_DEG = 5.0e-7 north (crates/cartopolis/src/systems/dev/shot_harness.rs:286, applied at :1071-1078). Each shot is placed through place_camera, which issues a full geo teleport (shot_harness.rs:2036). geo_nav::apply_teleport then calls stream.request_reanchor(lat, lng) unconditionally (crates/cartopolis/src/systems/nav/geo_nav.rs:2003), and map_stream::apply_reanchor (crates/cartopolis/src/systems/map/map_stream.rs:368) retires every loaded tile (:384-387), resets the fetch bookkeeping (:392), clears the quad meshes (:396), and bumps both anchor_generation (:406) and fetch_generation (:410) — without ever checking whether snap_anchor (:399) returned the anchor it already held.

For the flicker pair it always does. From job 1254's log, both captures:

09:26:00 camera placed shot=1 of=2  lat=53.2194    lng=6.5665
09:26:00 map re-anchored            lat=53.220013067568445 lng=6.565704345703125
09:34:17 camera placed shot=2 of=2  lat=53.2194005 lng=6.5665
09:34:17 map re-anchored            lat=53.220013067568445 lng=6.565704345703125   ← identical

and the second one is immediately followed by shaped the basemap revision=3, repeated lod22 cell rendered, parked vehicles hit the per-cell cap, moorings hit the per-cell cap and large GPU upload queued in one frame — a cold rebuild of a world the process already had in memory. Over twenty consumers key off anchor_generation and all of them retear with it (map_geometry.rs:942, lod22.rs:476, osm_buildings.rs:1016, street_detail.rs:1355, tall_structures.rs:588, interiors.rs:133, dikes.rs:341, building_facts.rs:678, terrain.rs:247, weather.rs:130, globe.rs:721, navigation.rs:1113, transit_vehicles.rs:239, data_layers.rs:986).

That second re-stream is ~28.5 virtual settle seconds, i.e. roughly half the step's wall clock. It is also not confined to CI: any in-app teleport landing inside the tile the map is already anchored on pays the same city-wide retear.

Why the cap cannot simply be raised

--wait and --settle accumulate the virtual clock — progress.elapsed += time.delta_secs() (shot_harness.rs:2400) and progress.quiet += time.delta_secs() (:2516) — which Bevy clamps at max_delta = 0.25 s, as the file's own comments state (:451, :1332, :2404-2406). On the CI runner a settled frame at this viewpoint costs ~2.5 s of wall clock (frame_ms="2551.0" in the failing job's shot state line), so one virtual second buys ~10 wall seconds. --wait 120 (ci.yml:469) therefore licenses ~20 minutes of wall clock per capture, and --flicker takes two. No timeout-minutes value bounds that; only making the work smaller does.

A secondary finding worth knowing

A timeout-minutes kill is enforced by the runner, so the step's own failure handler — || { echo "::error::…"; cat /tmp/flicker.json; exit 1; } (ci.yml:470-474) — never runs. The failing job's log contains zero ::error:: lines and no metrics dump; what it does contain is ~10,000 lines of platform::watchdog thread dumps. A timed-out smoke is currently undiagnosable from its own log.

Approach

One crate, one behavioural change plus its test.

crates/cartopolis/src/systems/map/map_stream.rs — apply_reanchor (:368). Compute snap_anchor(lat, lng, base_zoom) before touching anything. If the result equals the current stream.anchor, take the pending request and return, having retired nothing and bumped neither generation. Only when the snapped anchor differs (or stream.anchor is None) does the existing retear body run.

Two things the implementer must check while there:

  • The stream.tiles.reset() comment at :389-392 already anticipates the unchanged-anchor case — it exists so that tiles the retire loop just dropped are not treated as still in flight. Skipping the retire and the reset together is coherent; skipping only one is not.
  • published.geo (:404) and current_zoom (:405) are re-published on every re-anchor today. On the early-return path they are already correct by definition, but confirm rather than assume.

The teleport itself must still happen — geo_nav::apply_teleport places the avatar and camera after the request (geo_nav.rs:2009-2012) and that path is untouched.

Test. Add to the existing mod tests at map_stream.rs:1365: two request_reanchor calls whose lat/lng differ by less than one tile at the base zoom must leave anchor_generation unchanged after the second, while a call landing in a different tile must bump it. anchor_generation is pub (:189); fetch_generation is private (:200), so if the test needs to observe it, expose a reader rather than making the field public.

Do not raise timeout-minutes. With the second capture no longer re-streaming, 15 minutes becomes a real backstop rather than the thing being raced.

Acceptance criteria

  • apply_reanchor performs no tile retire, no tiles.reset(), no quads.clear(), and no anchor_generation / fetch_generation bump when the snapped anchor equals the one already held.
  • A re-anchor request that snaps to a different tile still does all of the above, exactly as today.
  • A unit test in map_stream.rs's mod tests pins both halves of the above.
  • A local --flicker run at the CI viewpoint logs map re-anchored once, not twice, and the second capture's log shows no shaped the basemap, no lod22 cell rendered and no parked vehicles hit the per-cell cap after camera placed shot=2.
  • flicker_px is re-read after the change and reported on this ticket with the measured number. The gate flicker_px<=1500 (ci.yml:470) is not loosened. If the measured floor moves away from the ~500 px recorded in ci.yml:449-458 (recent CI values: 169, 384, 461, 464, 510), say so and update that comment block with the new figure and its date — the old floor was measured on a second capture that was re-streaming, so part of it may have been arriving geometry rather than depth flicker.
  • cargo fmt --check passes workspace-wide.
  • The merge run's Headless flicker smoke (street) step completes well inside its 15-minute cap, and its duration is reported here.

Verification

# unit tests (the crate is a lib + shim bin, so this links the rlib)
cargo test -p cartopolis map_stream
cargo test -p cartopolis

# the exact CI command, minus the assertions, to read the log and the timing
CARTO_TILE_URL=https://tiles.cartopolis.org/tiles/osm/{z}/{x}/{y} \
WGPU_BACKEND=vulkan \
cargo run -p cartopolis -- \
  --shot /tmp/flicker.png --dump-state /tmp/flicker.json \
  --flicker --time 12 --fog 0 --clouds 0 --rain 0 \
  --at 53.2194,6.5665,100 --look=-60,0 \
  --size 640x360 --settle 2 --wait 120 --hide-ui

cargo fmt --check

CLAUDE.md's corrected note (2026‑08‑15) and the shot-harness-works-in-this-container memory say Mesa lavapipe is installed here, so the --flicker run above should work in this container. Two caveats: it renders on a CPU rasteriser, so colour, exposure and lighting are not evidence — geometry, placement and flicker_px are; and it will take several minutes even after the fix. If lavapipe is missing on the machine that picks this up, the two log-based criteria are checkable in CI instead, on the merge run.

The step-duration criterion can only be read from CI: after merge, fetch the job log and diff the timestamps between the Running \target/debug/cartopolis --shot /tmp/flicker.pngline andflicker pair measured`.

Out of scope

  • Raising timeout-minutes on any smoke step. The cap is not the defect.
  • Giving the shot harness a wall-clock --deadline. --wait being a virtual-second budget that cannot bound wall clock is real (shot_harness.rs:2400, :2516) and worth its own ticket, but it is a separate design change with its own trade-off — the file argues at :2404-2406 that the virtual clock is deliberate, since it measures the scripted run's intent.
  • Making a timeout-minutes kill diagnosable. The || { cat … } handler at ci.yml:470-474 cannot fire on a runner kill. Worth fixing; not this ticket.
  • The scene-smoke scope filter. 407c1e5a changed only tools/autopilot.ts, which is not on the deny-list at ci.yml:318, so it paid for three full renders. That is the documented fail-safe default ("the default for an unrecognised path … is to run the smokes"), and changing it is a policy decision, not a bug fix.
  • The watchdog's log volume. ~10,000 lines of platform::watchdog thread dumps in one job log is a real nuisance on a CPU rasteriser where 2.5 s frames are normal, but it is not why the step failed.

Open questions

None.


Branch: perf/237-reanchor-skips-an-unmoved-anchor

Original request

The last commit on main that CI ran for is 407c1e5a, and it did not pass:

  • test cartopolis — failure

Nothing can be deployed while this stands — tools/deploy-main gates on it — so this
comes before anything on the frontier.

Read the failing job's log, reproduce it locally, and fix the cause. If the failure is
the runner rather than the code (a flake, a cache miss, a missing tool), say so on this
ticket and close it rather than changing code to suit it.

Filed by the autopilot.

🤖 Refined by the viberfox issue agent. Reply with @agent refine and what is wrong to have this rewritten.

## Problem `main` is red at `407c1e5a` because the **`Headless flicker smoke (street)`** step (`.forgejo/workflows/ci.yml:459`) was killed by its own `timeout-minutes: 15` (`.forgejo/workflows/ci.yml:461`). Failing job: [run 682, job 1259](https://code.garage44.eu/jeroen/cartopolis/actions/runs/682). Nothing is wrong with the tests — every suite in that job reports `test result: ok`, 0 failed. The step started its `cargo run` at 09:49:33 and the runner logged `⚙️ [runner]: context deadline exceeded` at 10:04:18, 14 m 45 s later. It had saved capture 1 of the flicker pair and was 6 minutes into capture 2. **This is not a flake, and it is not the runner.** The step has been running at 60–95 % of its own cap since it landed (`4cebab4`, 2026‑08‑29). Every main run of it so far: | job | commit | capture 1 settle (virtual s) | mean frame ms | step wall clock | |---|---|---|---|---| | 1234 | `c815861d` | 30.1 | 1968 | 8 m 53 s | | 1242 | `20b036a1` | 45.6 | 2340 | 13 m 26 s | | 1246 | `ab430c5a` | 30.0 | 1984 | 9 m 04 s | | 1250 | `29405f29` | 45.7 | 2669 | 13 m 36 s | | 1254 | `bdd05d07` | 45.9 | 2557 | 13 m 31 s | | 1259 | `407c1e5a` | 45.8 | 2551 | **killed at 14 m 45 s** | ### Where the time actually goes **The second capture of a `--flicker` pair re-streams the entire city, for a 5 cm camera move.** `--flicker` turns each viewpoint into two shots, the twin nudged `FLICKER_NUDGE_DEG = 5.0e-7` north (`crates/cartopolis/src/systems/dev/shot_harness.rs:286`, applied at `:1071-1078`). Each shot is placed through `place_camera`, which issues a full geo teleport (`shot_harness.rs:2036`). `geo_nav::apply_teleport` then calls `stream.request_reanchor(lat, lng)` **unconditionally** (`crates/cartopolis/src/systems/nav/geo_nav.rs:2003`), and `map_stream::apply_reanchor` (`crates/cartopolis/src/systems/map/map_stream.rs:368`) retires every loaded tile (`:384-387`), resets the fetch bookkeeping (`:392`), clears the quad meshes (`:396`), and bumps both `anchor_generation` (`:406`) and `fetch_generation` (`:410`) — *without ever checking whether `snap_anchor` (`:399`) returned the anchor it already held.* For the flicker pair it always does. From job 1254's log, both captures: ``` 09:26:00 camera placed shot=1 of=2 lat=53.2194 lng=6.5665 09:26:00 map re-anchored lat=53.220013067568445 lng=6.565704345703125 09:34:17 camera placed shot=2 of=2 lat=53.2194005 lng=6.5665 09:34:17 map re-anchored lat=53.220013067568445 lng=6.565704345703125 ← identical ``` and the second one is immediately followed by `shaped the basemap revision=3`, repeated `lod22 cell rendered`, `parked vehicles hit the per-cell cap`, `moorings hit the per-cell cap` and `large GPU upload queued in one frame` — a cold rebuild of a world the process already had in memory. Over twenty consumers key off `anchor_generation` and all of them retear with it (`map_geometry.rs:942`, `lod22.rs:476`, `osm_buildings.rs:1016`, `street_detail.rs:1355`, `tall_structures.rs:588`, `interiors.rs:133`, `dikes.rs:341`, `building_facts.rs:678`, `terrain.rs:247`, `weather.rs:130`, `globe.rs:721`, `navigation.rs:1113`, `transit_vehicles.rs:239`, `data_layers.rs:986`). That second re-stream is ~28.5 virtual settle seconds, i.e. roughly half the step's wall clock. It is also not confined to CI: any in-app teleport landing inside the tile the map is already anchored on pays the same city-wide retear. ### Why the cap cannot simply be raised `--wait` and `--settle` accumulate the **virtual** clock — `progress.elapsed += time.delta_secs()` (`shot_harness.rs:2400`) and `progress.quiet += time.delta_secs()` (`:2516`) — which Bevy clamps at `max_delta` = 0.25 s, as the file's own comments state (`:451`, `:1332`, `:2404-2406`). On the CI runner a settled frame at this viewpoint costs ~2.5 s of wall clock (`frame_ms="2551.0"` in the failing job's `shot state` line), so one virtual second buys ~10 wall seconds. `--wait 120` (`ci.yml:469`) therefore licenses ~20 minutes of wall clock *per capture*, and `--flicker` takes two. No `timeout-minutes` value bounds that; only making the work smaller does. ### A secondary finding worth knowing A `timeout-minutes` kill is enforced by the runner, so the step's own failure handler — `|| { echo "::error::…"; cat /tmp/flicker.json; exit 1; }` (`ci.yml:470-474`) — never runs. The failing job's log contains **zero** `::error::` lines and no metrics dump; what it does contain is ~10,000 lines of `platform::watchdog` thread dumps. A timed-out smoke is currently undiagnosable from its own log. ## Approach One crate, one behavioural change plus its test. **`crates/cartopolis/src/systems/map/map_stream.rs` — `apply_reanchor` (`:368`).** Compute `snap_anchor(lat, lng, base_zoom)` *before* touching anything. If the result equals the current `stream.anchor`, take the pending request and return, having retired nothing and bumped neither generation. Only when the snapped anchor differs (or `stream.anchor` is `None`) does the existing retear body run. Two things the implementer must check while there: - The `stream.tiles.reset()` comment at `:389-392` already anticipates the unchanged-anchor case — it exists so that tiles the retire loop just dropped are not treated as still in flight. Skipping the retire and the reset together is coherent; skipping only one is not. - `published.geo` (`:404`) and `current_zoom` (`:405`) are re-published on every re-anchor today. On the early-return path they are already correct by definition, but confirm rather than assume. The teleport itself must still happen — `geo_nav::apply_teleport` places the avatar and camera after the request (`geo_nav.rs:2009-2012`) and that path is untouched. **Test.** Add to the existing `mod tests` at `map_stream.rs:1365`: two `request_reanchor` calls whose lat/lng differ by less than one tile at the base zoom must leave `anchor_generation` unchanged after the second, while a call landing in a different tile must bump it. `anchor_generation` is `pub` (`:189`); `fetch_generation` is private (`:200`), so if the test needs to observe it, expose a reader rather than making the field public. **Do not raise `timeout-minutes`.** With the second capture no longer re-streaming, 15 minutes becomes a real backstop rather than the thing being raced. ## Acceptance criteria - [ ] `apply_reanchor` performs no tile retire, no `tiles.reset()`, no `quads.clear()`, and no `anchor_generation` / `fetch_generation` bump when the snapped anchor equals the one already held. - [ ] A re-anchor request that snaps to a *different* tile still does all of the above, exactly as today. - [ ] A unit test in `map_stream.rs`'s `mod tests` pins both halves of the above. - [ ] A local `--flicker` run at the CI viewpoint logs `map re-anchored` **once**, not twice, and the second capture's log shows no `shaped the basemap`, no `lod22 cell rendered` and no `parked vehicles hit the per-cell cap` after `camera placed shot=2`. - [ ] `flicker_px` is re-read after the change and reported on this ticket with the measured number. The gate `flicker_px<=1500` (`ci.yml:470`) is **not** loosened. If the measured floor moves away from the ~500 px recorded in `ci.yml:449-458` (recent CI values: 169, 384, 461, 464, 510), say so and update that comment block with the new figure and its date — the old floor was measured on a second capture that was re-streaming, so part of it may have been arriving geometry rather than depth flicker. - [ ] `cargo fmt --check` passes workspace-wide. - [ ] The merge run's `Headless flicker smoke (street)` step completes well inside its 15-minute cap, and its duration is reported here. ## Verification ```bash # unit tests (the crate is a lib + shim bin, so this links the rlib) cargo test -p cartopolis map_stream cargo test -p cartopolis # the exact CI command, minus the assertions, to read the log and the timing CARTO_TILE_URL=https://tiles.cartopolis.org/tiles/osm/{z}/{x}/{y} \ WGPU_BACKEND=vulkan \ cargo run -p cartopolis -- \ --shot /tmp/flicker.png --dump-state /tmp/flicker.json \ --flicker --time 12 --fog 0 --clouds 0 --rain 0 \ --at 53.2194,6.5665,100 --look=-60,0 \ --size 640x360 --settle 2 --wait 120 --hide-ui cargo fmt --check ``` `CLAUDE.md`'s corrected note (2026‑08‑15) and the `shot-harness-works-in-this-container` memory say Mesa lavapipe is installed here, so the `--flicker` run above should work in this container. Two caveats: it renders on a CPU rasteriser, so **colour, exposure and lighting are not evidence** — geometry, placement and `flicker_px` are; and it will take several minutes even after the fix. If lavapipe is missing on the machine that picks this up, the two log-based criteria are checkable in CI instead, on the merge run. The step-duration criterion can only be read from CI: after merge, fetch the job log and diff the timestamps between the `Running \`target/debug/cartopolis --shot /tmp/flicker.png` line and `flicker pair measured`. ## Out of scope - **Raising `timeout-minutes` on any smoke step.** The cap is not the defect. - **Giving the shot harness a wall-clock `--deadline`.** `--wait` being a virtual-second budget that cannot bound wall clock is real (`shot_harness.rs:2400`, `:2516`) and worth its own ticket, but it is a separate design change with its own trade-off — the file argues at `:2404-2406` that the virtual clock is deliberate, since it measures the scripted run's intent. - **Making a `timeout-minutes` kill diagnosable.** The `|| { cat … }` handler at `ci.yml:470-474` cannot fire on a runner kill. Worth fixing; not this ticket. - **The scene-smoke scope filter.** `407c1e5a` changed only `tools/autopilot.ts`, which is not on the deny-list at `ci.yml:318`, so it paid for three full renders. That is the documented fail-safe default ("the default for an unrecognised path … is to run the smokes"), and changing it is a policy decision, not a bug fix. - **The watchdog's log volume.** ~10,000 lines of `platform::watchdog` thread dumps in one job log is a real nuisance on a CPU rasteriser where 2.5 s frames are normal, but it is not why the step failed. ## Open questions None. --- Branch: `perf/237-reanchor-skips-an-unmoved-anchor` <details><summary>Original request</summary> The last commit on `main` that CI ran for is `407c1e5a`, and it did not pass: - `test cartopolis` — failure Nothing can be deployed while this stands — `tools/deploy-main` gates on it — so this comes before anything on the frontier. Read the failing job's log, reproduce it locally, and fix the cause. If the failure is the runner rather than the code (a flake, a cache miss, a missing tool), say so on this ticket and close it rather than changing code to suit it. <sub>Filed by the autopilot. <!-- autopilot:main-red-407c1e5a --></sub> </details> <sub>🤖 Refined by the viberfox issue agent. Reply with **@agent refine** and what is wrong to have this rewritten.</sub>
Author
Collaborator

🤖 Refining this before anything gets built — reading the code, no branch, no commits. If it turns out I'd have to guess at the requirements I'll stop and ask instead.

🤖 Refining this before anything gets built — reading the code, no branch, no commits. If it turns out I'd have to guess at the requirements I'll stop and ask instead.
Author
Collaborator

🤖 Refined. The description above is now the spec; I'll pick it up for implementation on the next poll.

If the direction is wrong, say so now — @agent refine with a correction rewrites it, and agent:skip stops me.


🤖 viberfox issue agent on the VPS — 60 turns · $5.22 · session 58fe7b8f-6825-4c4b-999b-dd72420e907c. Reply with @agent to continue this session.

🤖 **Refined.** The description above is now the spec; I'll pick it up for implementation on the next poll. If the direction is wrong, say so now — **@agent refine** with a correction rewrites it, and `agent:skip` stops me. --- <sub>🤖 viberfox issue agent on the VPS — 60 turns · $5.22 · session `58fe7b8f-6825-4c4b-999b-dd72420e907c`. Reply with **@agent** to continue this session.</sub>
Author
Collaborator

🤖 Picked this up — working on branch perf/237-reanchor-skips-an-unmoved-anchor in a fresh worktree on the VPS. I'll report back here when I'm done.

Label an issue agent:skip to keep me off it, or driven to keep it in a session you are driving.

🤖 Picked this up — working on branch `perf/237-reanchor-skips-an-unmoved-anchor` in a fresh worktree on the VPS. I'll report back here when I'm done. <sub>Label an issue `agent:skip` to keep me off it, or `driven` to keep it in a session you are driving.</sub>
Author
Collaborator

apply_reanchor in crates/cartopolis/src/systems/map/map_stream.rs now checks where a re-anchor request actually lands before it tears anything down. If the request snaps to the anchor the map already holds, it consumes the request and returns — no tiles retired, no fetch bookkeeping reset, no quad meshes cleared, neither generation bumped. A request that snaps to a different tile does exactly what it did before.

That was the whole cost of the timed-out step. The map anchor is a tile centre, so a --flicker pair's 5 cm nudge always snapped to the same one, and the second capture rebuilt Groningen from scratch.

What I measured

Both runs are the exact CI command, on this container's software renderer, same viewpoint and same weather/time pins. One sample each.

map re-anchored lines capture 2 settle capture 2 wall clock flicker_px
before 2 28.5 virtual seconds 4 min 20 s 553
after 1 2.0 virtual seconds 20 s 714

The second capture's log after the fix has no shaped the basemap, no lod22 cell rendered and no parked vehicles hit the per-cell cap — all four acceptance criteria for the log. Total placement-to-measurement was 6 min 15 s, and the remaining 5 min 54 s of that is the first capture's settle, which this change does not touch.

I also read back the second capture's image: a fully streamed street scene with buildings, trees, pavement and a street label. Skipping the teardown did not leave it with a half-built world.

The number that did not go the way the ticket expected

flicker_px went up, 553 → 714, not down. The ticket's guess was that the old floor was inflated by geometry still arriving during the re-stream; on this evidence it was not. But the first six runs on main already gave 169, 384, 461, 464 and 510 for the same shot — a threefold spread — so two samples cannot tell "the fix raised the floor" apart from "this metric is noisy". I did not loosen the gate; it stays at flicker_px<=1500, and 714 sits well under it.

I rewrote the calibration comment above that step in .forgejo/workflows/ci.yml to record both figures with today's date, to say the floor is noisy, and to say to read a trend across runs rather than one number. That file change is comments only.

Verified

  • cargo test -p cartopolis — 815 passed, 0 failed, including the new test.
  • cargo test -p cartopolis_geo — 236 passed, 0 failed.
  • cargo fmt --check across the workspace — clean.

The new test spawns a tile, asks for a re-anchor 5 cm away, and asserts neither generation moved, the tile survived, the published anchor still agrees with the streamer's, and the request was consumed — then asks for one several tiles east and asserts the full teardown still happens.

For you to decide

  • Is the raised flicker floor worth a look? I would not block on it — it is one sample against a metric with a threefold spread, and 714 is less than half the gate. But if you want certainty, the cheapest answer is to let the next few runs on main accumulate and see where it settles.
  • The viewpoint sits inside one tile by luck. If anyone ever moves the flicker shot to a spot where the 5 cm nudge crosses a tile boundary, the second capture will re-stream again and the step will go back to racing its cap. Nothing in the code prevents that. It did not seem worth a guard for a single hardcoded coordinate, but it is a trap worth knowing about.
  • Two things the ticket listed as out of scope are still real: --wait is a budget in virtual seconds and cannot bound wall clock, and a timeout-minutes kill silently skips the step's own cat /tmp/flicker.json handler, so a timed-out run still explains nothing about itself.

Branch perf/237-reanchor-skips-an-unmoved-anchor is pushed with two commits. The step's duration on the merge run still has to be read from CI — I can only report what this container measured.

Branch perf/237-reanchor-skips-an-unmoved-anchor (pushed)

c0e8fa6 docs(ci): re-measure the flicker floor after the re-anchor fix
c5ddf41 perf(map): a re-anchor that does not move the anchor costs nothing

🤖 viberfox issue agent on the VPS — 173 turns · $14.36 · session fb8e516e-e0eb-453d-ba35-197e8f15f4ee. Reply with @agent to continue this session.

`apply_reanchor` in `crates/cartopolis/src/systems/map/map_stream.rs` now checks where a re-anchor request actually lands **before** it tears anything down. If the request snaps to the anchor the map already holds, it consumes the request and returns — no tiles retired, no fetch bookkeeping reset, no quad meshes cleared, neither generation bumped. A request that snaps to a different tile does exactly what it did before. That was the whole cost of the timed-out step. The map anchor is a tile centre, so a `--flicker` pair's 5 cm nudge always snapped to the same one, and the second capture rebuilt Groningen from scratch. ## What I measured Both runs are the exact CI command, on this container's software renderer, same viewpoint and same weather/time pins. One sample each. | | `map re-anchored` lines | capture 2 settle | capture 2 wall clock | `flicker_px` | |---|---|---|---|---| | before | 2 | 28.5 virtual seconds | 4 min 20 s | 553 | | after | **1** | **2.0 virtual seconds** | **20 s** | **714** | The second capture's log after the fix has no `shaped the basemap`, no `lod22 cell rendered` and no `parked vehicles hit the per-cell cap` — all four acceptance criteria for the log. Total placement-to-measurement was 6 min 15 s, and the remaining 5 min 54 s of that is the *first* capture's settle, which this change does not touch. I also read back the second capture's image: a fully streamed street scene with buildings, trees, pavement and a street label. Skipping the teardown did not leave it with a half-built world. ## The number that did not go the way the ticket expected `flicker_px` went **up**, 553 → 714, not down. The ticket's guess was that the old floor was inflated by geometry still arriving during the re-stream; on this evidence it was not. But the first six runs on `main` already gave 169, 384, 461, 464 and 510 for the same shot — a threefold spread — so two samples cannot tell "the fix raised the floor" apart from "this metric is noisy". I did not loosen the gate; it stays at `flicker_px<=1500`, and 714 sits well under it. I rewrote the calibration comment above that step in `.forgejo/workflows/ci.yml` to record both figures with today's date, to say the floor is noisy, and to say to read a trend across runs rather than one number. That file change is comments only. ## Verified - `cargo test -p cartopolis` — 815 passed, 0 failed, including the new test. - `cargo test -p cartopolis_geo` — 236 passed, 0 failed. - `cargo fmt --check` across the workspace — clean. The new test spawns a tile, asks for a re-anchor 5 cm away, and asserts neither generation moved, the tile survived, the published anchor still agrees with the streamer's, and the request was consumed — then asks for one several tiles east and asserts the full teardown still happens. ## For you to decide - **Is the raised flicker floor worth a look?** I would not block on it — it is one sample against a metric with a threefold spread, and 714 is less than half the gate. But if you want certainty, the cheapest answer is to let the next few runs on `main` accumulate and see where it settles. - **The viewpoint sits inside one tile by luck.** If anyone ever moves the flicker shot to a spot where the 5 cm nudge crosses a tile boundary, the second capture will re-stream again and the step will go back to racing its cap. Nothing in the code prevents that. It did not seem worth a guard for a single hardcoded coordinate, but it is a trap worth knowing about. - Two things the ticket listed as out of scope are still real: `--wait` is a budget in virtual seconds and cannot bound wall clock, and a `timeout-minutes` kill silently skips the step's own `cat /tmp/flicker.json` handler, so a timed-out run still explains nothing about itself. Branch `perf/237-reanchor-skips-an-unmoved-anchor` is pushed with two commits. The step's duration on the merge run still has to be read from CI — I can only report what this container measured. **Branch** `perf/237-reanchor-skips-an-unmoved-anchor` (pushed) ``` c0e8fa6 docs(ci): re-measure the flicker floor after the re-anchor fix c5ddf41 perf(map): a re-anchor that does not move the anchor costs nothing ``` --- <sub>🤖 viberfox issue agent on the VPS — 173 turns · $14.36 · session `fb8e516e-e0eb-453d-ba35-197e8f15f4ee`. Reply with **@agent** to continue this session.</sub>
Author
Collaborator

🤖 Marked agent:done without a CI verdict on c0e8fa66 — no workflow reported on this commit. The branch is waiting for review; check the run yourself before landing it.

🤖 Marked `agent:done` **without a CI verdict** on `c0e8fa66` — no workflow reported on this commit. The branch is waiting for review; check the run yourself before landing it.
Author
Collaborator

🤖 Not shipping this one — it touches something the unattended lane may not change:

  • .forgejo/workflows/ci.yml — the CI or deploy workflows

The branch perf/237-reanchor-skips-an-unmoved-anchor is finished and CI passed on it; it is waiting for you to land it. The reasoning is in docs/direction.md — these are the changes no revert from here undoes.

🤖 **Not shipping this one — it touches something the unattended lane may not change:** - `.forgejo/workflows/ci.yml` — the CI or deploy workflows The branch `perf/237-reanchor-skips-an-unmoved-anchor` is finished and CI passed on it; it is waiting for you to land it. The reasoning is in `docs/direction.md` — these are the changes no revert from here undoes.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
jeroen/cartopolis#237
No description provided.