main is red on 407c1e5a #237
Labels
No labels
agent
agent:ci
agent:done
agent:failed
agent:needs-input
agent:refined
agent:refining
agent:running
agent:shipped
agent:skip
autonomous
autopilot
driven
local
plan
proposal
qa
qa-gap
research
retro
ship
No milestone
No project
No assignees
1 participant
Notifications
Due date
No due date set.
Dependencies
No dependencies set
Reference
jeroen/cartopolis#237
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Problem
mainis red at407c1e5abecause theHeadless flicker smoke (street)step (.forgejo/workflows/ci.yml:459) was killed by its owntimeout-minutes: 15(.forgejo/workflows/ci.yml:461). Failing job: run 682, job 1259.Nothing is wrong with the tests — every suite in that job reports
test result: ok, 0 failed. The step started itscargo runat 09:49:33 and the runner logged⚙️ [runner]: context deadline exceededat 10:04:18, 14 m 45 s later. It had saved capture 1 of the flicker pair and was 6 minutes into capture 2.This is not a flake, and it is not the runner. The step has been running at 60–95 % of its own cap since it landed (
4cebab4, 2026‑08‑29). Every main run of it so far:c815861d20b036a1ab430c5a29405f29bdd05d07407c1e5aWhere the time actually goes
The second capture of a
--flickerpair re-streams the entire city, for a 5 cm camera move.--flickerturns each viewpoint into two shots, the twin nudgedFLICKER_NUDGE_DEG = 5.0e-7north (crates/cartopolis/src/systems/dev/shot_harness.rs:286, applied at:1071-1078). Each shot is placed throughplace_camera, which issues a full geo teleport (shot_harness.rs:2036).geo_nav::apply_teleportthen callsstream.request_reanchor(lat, lng)unconditionally (crates/cartopolis/src/systems/nav/geo_nav.rs:2003), andmap_stream::apply_reanchor(crates/cartopolis/src/systems/map/map_stream.rs:368) retires every loaded tile (:384-387), resets the fetch bookkeeping (:392), clears the quad meshes (:396), and bumps bothanchor_generation(:406) andfetch_generation(:410) — without ever checking whethersnap_anchor(:399) returned the anchor it already held.For the flicker pair it always does. From job 1254's log, both captures:
and the second one is immediately followed by
shaped the basemap revision=3, repeatedlod22 cell rendered,parked vehicles hit the per-cell cap,moorings hit the per-cell capandlarge GPU upload queued in one frame— a cold rebuild of a world the process already had in memory. Over twenty consumers key offanchor_generationand all of them retear with it (map_geometry.rs:942,lod22.rs:476,osm_buildings.rs:1016,street_detail.rs:1355,tall_structures.rs:588,interiors.rs:133,dikes.rs:341,building_facts.rs:678,terrain.rs:247,weather.rs:130,globe.rs:721,navigation.rs:1113,transit_vehicles.rs:239,data_layers.rs:986).That second re-stream is ~28.5 virtual settle seconds, i.e. roughly half the step's wall clock. It is also not confined to CI: any in-app teleport landing inside the tile the map is already anchored on pays the same city-wide retear.
Why the cap cannot simply be raised
--waitand--settleaccumulate the virtual clock —progress.elapsed += time.delta_secs()(shot_harness.rs:2400) andprogress.quiet += time.delta_secs()(:2516) — which Bevy clamps atmax_delta= 0.25 s, as the file's own comments state (:451,:1332,:2404-2406). On the CI runner a settled frame at this viewpoint costs ~2.5 s of wall clock (frame_ms="2551.0"in the failing job'sshot stateline), so one virtual second buys ~10 wall seconds.--wait 120(ci.yml:469) therefore licenses ~20 minutes of wall clock per capture, and--flickertakes two. Notimeout-minutesvalue bounds that; only making the work smaller does.A secondary finding worth knowing
A
timeout-minuteskill is enforced by the runner, so the step's own failure handler —|| { echo "::error::…"; cat /tmp/flicker.json; exit 1; }(ci.yml:470-474) — never runs. The failing job's log contains zero::error::lines and no metrics dump; what it does contain is ~10,000 lines ofplatform::watchdogthread dumps. A timed-out smoke is currently undiagnosable from its own log.Approach
One crate, one behavioural change plus its test.
crates/cartopolis/src/systems/map/map_stream.rs—apply_reanchor(:368). Computesnap_anchor(lat, lng, base_zoom)before touching anything. If the result equals the currentstream.anchor, take the pending request and return, having retired nothing and bumped neither generation. Only when the snapped anchor differs (orstream.anchorisNone) does the existing retear body run.Two things the implementer must check while there:
stream.tiles.reset()comment at:389-392already anticipates the unchanged-anchor case — it exists so that tiles the retire loop just dropped are not treated as still in flight. Skipping the retire and the reset together is coherent; skipping only one is not.published.geo(:404) andcurrent_zoom(:405) are re-published on every re-anchor today. On the early-return path they are already correct by definition, but confirm rather than assume.The teleport itself must still happen —
geo_nav::apply_teleportplaces the avatar and camera after the request (geo_nav.rs:2009-2012) and that path is untouched.Test. Add to the existing
mod testsatmap_stream.rs:1365: tworequest_reanchorcalls whose lat/lng differ by less than one tile at the base zoom must leaveanchor_generationunchanged after the second, while a call landing in a different tile must bump it.anchor_generationispub(:189);fetch_generationis private (:200), so if the test needs to observe it, expose a reader rather than making the field public.Do not raise
timeout-minutes. With the second capture no longer re-streaming, 15 minutes becomes a real backstop rather than the thing being raced.Acceptance criteria
apply_reanchorperforms no tile retire, notiles.reset(), noquads.clear(), and noanchor_generation/fetch_generationbump when the snapped anchor equals the one already held.map_stream.rs'smod testspins both halves of the above.--flickerrun at the CI viewpoint logsmap re-anchoredonce, not twice, and the second capture's log shows noshaped the basemap, nolod22 cell renderedand noparked vehicles hit the per-cell capaftercamera placed shot=2.flicker_pxis re-read after the change and reported on this ticket with the measured number. The gateflicker_px<=1500(ci.yml:470) is not loosened. If the measured floor moves away from the ~500 px recorded inci.yml:449-458(recent CI values: 169, 384, 461, 464, 510), say so and update that comment block with the new figure and its date — the old floor was measured on a second capture that was re-streaming, so part of it may have been arriving geometry rather than depth flicker.cargo fmt --checkpasses workspace-wide.Headless flicker smoke (street)step completes well inside its 15-minute cap, and its duration is reported here.Verification
CLAUDE.md's corrected note (2026‑08‑15) and theshot-harness-works-in-this-containermemory say Mesa lavapipe is installed here, so the--flickerrun above should work in this container. Two caveats: it renders on a CPU rasteriser, so colour, exposure and lighting are not evidence — geometry, placement andflicker_pxare; and it will take several minutes even after the fix. If lavapipe is missing on the machine that picks this up, the two log-based criteria are checkable in CI instead, on the merge run.The step-duration criterion can only be read from CI: after merge, fetch the job log and diff the timestamps between the
Running \target/debug/cartopolis --shot /tmp/flicker.pngline andflicker pair measured`.Out of scope
timeout-minuteson any smoke step. The cap is not the defect.--deadline.--waitbeing a virtual-second budget that cannot bound wall clock is real (shot_harness.rs:2400,:2516) and worth its own ticket, but it is a separate design change with its own trade-off — the file argues at:2404-2406that the virtual clock is deliberate, since it measures the scripted run's intent.timeout-minuteskill diagnosable. The|| { cat … }handler atci.yml:470-474cannot fire on a runner kill. Worth fixing; not this ticket.407c1e5achanged onlytools/autopilot.ts, which is not on the deny-list atci.yml:318, so it paid for three full renders. That is the documented fail-safe default ("the default for an unrecognised path … is to run the smokes"), and changing it is a policy decision, not a bug fix.platform::watchdogthread dumps in one job log is a real nuisance on a CPU rasteriser where 2.5 s frames are normal, but it is not why the step failed.Open questions
None.
Branch:
perf/237-reanchor-skips-an-unmoved-anchorOriginal request
The last commit on
mainthat CI ran for is407c1e5a, and it did not pass:test cartopolis— failureNothing can be deployed while this stands —
tools/deploy-maingates on it — so thiscomes before anything on the frontier.
Read the failing job's log, reproduce it locally, and fix the cause. If the failure is
the runner rather than the code (a flake, a cache miss, a missing tool), say so on this
ticket and close it rather than changing code to suit it.
Filed by the autopilot.
🤖 Refined by the viberfox issue agent. Reply with @agent refine and what is wrong to have this rewritten.
🤖 Refining this before anything gets built — reading the code, no branch, no commits. If it turns out I'd have to guess at the requirements I'll stop and ask instead.
🤖 Refined. The description above is now the spec; I'll pick it up for implementation on the next poll.
If the direction is wrong, say so now — @agent refine with a correction rewrites it, and
agent:skipstops me.🤖 viberfox issue agent on the VPS — 60 turns · $5.22 · session
58fe7b8f-6825-4c4b-999b-dd72420e907c. Reply with @agent to continue this session.🤖 Picked this up — working on branch
perf/237-reanchor-skips-an-unmoved-anchorin a fresh worktree on the VPS. I'll report back here when I'm done.Label an issue
agent:skipto keep me off it, ordrivento keep it in a session you are driving.apply_reanchorincrates/cartopolis/src/systems/map/map_stream.rsnow checks where a re-anchor request actually lands before it tears anything down. If the request snaps to the anchor the map already holds, it consumes the request and returns — no tiles retired, no fetch bookkeeping reset, no quad meshes cleared, neither generation bumped. A request that snaps to a different tile does exactly what it did before.That was the whole cost of the timed-out step. The map anchor is a tile centre, so a
--flickerpair's 5 cm nudge always snapped to the same one, and the second capture rebuilt Groningen from scratch.What I measured
Both runs are the exact CI command, on this container's software renderer, same viewpoint and same weather/time pins. One sample each.
map re-anchoredlinesflicker_pxThe second capture's log after the fix has no
shaped the basemap, nolod22 cell renderedand noparked vehicles hit the per-cell cap— all four acceptance criteria for the log. Total placement-to-measurement was 6 min 15 s, and the remaining 5 min 54 s of that is the first capture's settle, which this change does not touch.I also read back the second capture's image: a fully streamed street scene with buildings, trees, pavement and a street label. Skipping the teardown did not leave it with a half-built world.
The number that did not go the way the ticket expected
flicker_pxwent up, 553 → 714, not down. The ticket's guess was that the old floor was inflated by geometry still arriving during the re-stream; on this evidence it was not. But the first six runs onmainalready gave 169, 384, 461, 464 and 510 for the same shot — a threefold spread — so two samples cannot tell "the fix raised the floor" apart from "this metric is noisy". I did not loosen the gate; it stays atflicker_px<=1500, and 714 sits well under it.I rewrote the calibration comment above that step in
.forgejo/workflows/ci.ymlto record both figures with today's date, to say the floor is noisy, and to say to read a trend across runs rather than one number. That file change is comments only.Verified
cargo test -p cartopolis— 815 passed, 0 failed, including the new test.cargo test -p cartopolis_geo— 236 passed, 0 failed.cargo fmt --checkacross the workspace — clean.The new test spawns a tile, asks for a re-anchor 5 cm away, and asserts neither generation moved, the tile survived, the published anchor still agrees with the streamer's, and the request was consumed — then asks for one several tiles east and asserts the full teardown still happens.
For you to decide
mainaccumulate and see where it settles.--waitis a budget in virtual seconds and cannot bound wall clock, and atimeout-minuteskill silently skips the step's owncat /tmp/flicker.jsonhandler, so a timed-out run still explains nothing about itself.Branch
perf/237-reanchor-skips-an-unmoved-anchoris pushed with two commits. The step's duration on the merge run still has to be read from CI — I can only report what this container measured.Branch
perf/237-reanchor-skips-an-unmoved-anchor(pushed)🤖 viberfox issue agent on the VPS — 173 turns · $14.36 · session
fb8e516e-e0eb-453d-ba35-197e8f15f4ee. Reply with @agent to continue this session.🤖 Marked
agent:donewithout a CI verdict onc0e8fa66— no workflow reported on this commit. The branch is waiting for review; check the run yourself before landing it.🤖 Not shipping this one — it touches something the unattended lane may not change:
.forgejo/workflows/ci.yml— the CI or deploy workflowsThe branch
perf/237-reanchor-skips-an-unmoved-anchoris finished and CI passed on it; it is waiting for you to land it. The reasoning is indocs/direction.md— these are the changes no revert from here undoes.Stuck, waiting for you: nonewhile three tickets wait for a person #270alreadyFiledonly ever sees 50 issues, so every autopilot dedupe key silently expires — #206 was QA'd twice #281alreadyFiledonly ever sees 50 issues, so every autopilot dedupe key silently expires — #206 was QA'd twice #281alreadyFiledonly ever sees 50 issues, so every autopilot dedupe key silently expires — #206 was QA'd twice #281alreadyFiledonly ever sees 50 issues, so every autopilot dedupe key silently expires — #206 was QA'd twice #281alreadyFiledonly ever sees 50 issues, so every autopilot dedupe key silently expires — #206 was QA'd twice #281alreadyFiledonly ever sees 50 issues, so every autopilot dedupe key silently expires — #206 was QA'd twice #281