Plan the focus: Improving performance on the steamdeck #292

Closed
opened 2026-09-09 21:29:01 +00:00 by viberfox-agent · 2 comments
Collaborator

The maintainer has set a focus for this lane, and it needs turning into a queue of work.

Improving performance on the steamdeck

This is a planning pass, not a build. Change no code and open no branch. What comes
back is a list, which the maintainer reorders and edits in the console before any of it
is filed — so propose the work, do not start it.

  1. Read before proposing. docs/direction.md for the goal and the seven values,
    CLAUDE.md for how this tree is built and what it already does, and whatever part of
    the code the focus names. Half of a good plan is noticing the thing is already there.
  2. Propose six to ten pieces of work. Each one a ticket a session could finish in a
    sitting, ordered so the earlier ones unblock the later ones. A title that says what
    changes, and a note saying where in the code it lands and how a machine would judge
    it — that last part is what this lane can land unattended.
  3. Leave out what it may not do. The wire protocol, database migrations, the CI
    workflows, tools/autopilot.ts itself and assetlinks.json are fenced: a ticket
    touching one is built and then parked for a person. Anything whose success is a
    matter of taste belongs to a person too.
  4. Post the plan in one comment on this ticket, as exactly this block:
```autopilot-plan
[{"title": "…", "note": "…"}, {"title": "…", "note": "…"}]
```

JSON, an array, title and note both strings. The autopilot reads the last such
block on this ticket and stores it as the queue; everything else in the comment is
ignored, so write whatever reasoning you want around it.
5. Then close this ticket. There is no branch to merge.

Decide it yourself. Nobody is watching this ticket, so a question asked here is a
ticket that stops. The values at the top of docs/direction.md settle the ambiguity, and
a piece of work you are unsure about is one to leave off the list rather than ask about.

Filed by the autopilot.

The maintainer has set a focus for this lane, and it needs turning into a queue of work. > Improving performance on the steamdeck **This is a planning pass, not a build.** Change no code and open no branch. What comes back is a list, which the maintainer reorders and edits in the console before any of it is filed — so propose the work, do not start it. 1. **Read before proposing.** `docs/direction.md` for the goal and the seven values, `CLAUDE.md` for how this tree is built and what it already does, and whatever part of the code the focus names. Half of a good plan is noticing the thing is already there. 2. **Propose six to ten pieces of work.** Each one a ticket a session could finish in a sitting, ordered so the earlier ones unblock the later ones. A title that says what changes, and a note saying where in the code it lands and how a machine would judge it — that last part is what this lane can land unattended. 3. **Leave out what it may not do.** The wire protocol, database migrations, the CI workflows, `tools/autopilot.ts` itself and `assetlinks.json` are fenced: a ticket touching one is built and then parked for a person. Anything whose success is a matter of taste belongs to a person too. 4. **Post the plan in one comment on this ticket**, as exactly this block: ```` ```autopilot-plan [{"title": "…", "note": "…"}, {"title": "…", "note": "…"}] ``` ```` JSON, an array, `title` and `note` both strings. The autopilot reads the last such block on this ticket and stores it as the queue; everything else in the comment is ignored, so write whatever reasoning you want around it. 5. **Then close this ticket.** There is no branch to merge. **Decide it yourself.** Nobody is watching this ticket, so a question asked here is a ticket that stops. The values at the top of `docs/direction.md` settle the ambiguity, and a piece of work you are unsure about is one to leave off the list rather than ask about. <sub>Filed by the autopilot. <!-- autopilot:focus-plan-037baf7d-5 --></sub>
Author
Collaborator

🤖 Picked this up — working on branch perf/292-plan-focus-improving-performance in a fresh worktree on the VPS. I'll report back here when I'm done.

Label an issue agent:skip to keep me off it, or driven to keep it in a session you are driving.

🤖 Picked this up — working on branch `perf/292-plan-focus-improving-performance` in a fresh worktree on the VPS. I'll report back here when I'm done. <sub>Label an issue `agent:skip` to keep me off it, or `driven` to keep it in a session you are driving.</sub>
Author
Collaborator

No code changed and the branch has no commits — this was a planning pass. I read docs/direction.md, CLAUDE.md, quality.rs, lib.rs's task-pool setup, frame_pacing.rs, map_stream.rs and the performance, streaming and allocation-budget notes.

The main finding: a Steam Deck is native Linux, so quality::auto() hands it Budget::desktop() — four shadow cascades to 3 km, 2× multisampling, 512-pixel tile rasters, the sky lighting cube re-baked every frame, and no ceiling at all on either upload rate or how many bytes of streamed map stay resident. systems::load_control therefore never runs on it. The rung that already describes the machine in its own documentation is Budget::balanced — "an integrated GPU, a thin laptop" — and nothing in the tree can currently reach it on a Deck.

Four things I checked that are already done, so the queue leaves them alone: the idle frame-rate throttle and the sleep limiter both work on the Linux path; target-cpu=x86-64-v3 already covers the Deck's Zen 2 cores; coarse basemap tiles fully covered by finer ones are already culled (map_stream::occluded_by_finer); and the tile pipeline is already gated on allocation counts.

Left off deliberately: the globe cloud shader, which the survey says is unmeasurable on this container's software renderer and needs the device; anything about colour, lighting or how the picture reads; and controller input, which is not performance.

Every number in these tickets has to be measured headlessly at 1280×800 here. Nobody in this lane has a Steam Deck, so confirmation on the actual hardware is yours.

[{"title": "Add a handheld render rung and reach it with CARTO_QUALITY=steamdeck", "note": "Lands in crates/cartopolis/src/quality.rs. Add handheld_guarded(Budget::balanced()) in the same shape as android_budget/handset_guarded, with the upload ceiling and resident-bytes target measured rather than guessed, and add the words steamdeck and handheld to parse(). No new QualityPreset variant, so the settings picker and user.json are untouched and the rung is reached by name or by auto-detection. Judged by cargo test -p cartopolis (the existing rung-ladder and tier-vocabulary tests, extended) and by a --shot --size 1280x800 --dump-state pair over Groningen at 150 m and 1700 m: draws, tris_shadow and gpu_resident_kb all below the desktop rung at the same viewpoint, with settled still true. Write the measured table into docs/notes/steam-deck.md the way handset_guarded's table is written."}, {"title": "Answer the handheld rung automatically on a Steam Deck", "note": "quality::auto()'s Linux arm and quality::on_this_machine, which today applies its guards only on Android. quality::init runs before a render adapter exists (the worker pool is sized from its answer), so the adapter cannot be asked; the facts available that early are the machine naming itself: SteamDeck=1 in the environment, and /sys/class/dmi/id/product_name of Jupiter or Galileo with board_vendor Valve. Keep the detection as a pure function over those three strings so it is testable off the device, and route it through on_this_machine so a rung picked by hand in Settings cannot delete a ceiling the machine needs, exactly as the handset does. Judged by unit tests covering both Deck models, a non-Valve Linux machine, and a missing /sys, which must still answer desktop; plus report_budget's chose= field naming the detector in the log."}, {"title": "Size the task pools and the tile-worker pool for four physical cores", "note": "lib.rs, both the TaskPoolThreadAssignmentPolicy block and the platform::workers::start call above it. A Steam Deck is four cores with hyperthreading, so available_parallelism answers 8 and the tile-worker pool takes 7: seven blocking rasterise jobs of roughly 120 ms each, alongside 4 io, 4 async-compute and 3 compute threads, on four physical cores sharing one 15 watt package with the GPU. This is the same arithmetic that silently left Android's parallel system executor one thread wide, and the survey's own lesson is that it was invisible because the widths had only ever been checked on a 22-core machine. Extract the split into a plain function the way frame_pacing::resolve_frame_rate was extracted, and set worker_cap on the handheld rung. Judged by unit tests of that function at 4, 8, 16 and 22 reported threads, and by settled_s and worker_outstanding at a fixed viewpoint not regressing in --dump-state."}, {"title": "Stop three thread::sleep calls from holding io-pool threads in osm_buildings", "note": "crates/cartopolis/src/systems/map/osm_buildings.rs: three paths inside async functions call std::thread::sleep, so each parks an io-pool thread instead of yielding. docs/notes/render-performance-survey.md lists this as open and rejects adding async-io as a direct dependency, but that was written before platform::fetch_gate and the Dispatcher landed on 2026-09-08, which give somewhere to re-rank the work and retry on a later frame instead of waiting on the thread. The io pool floors at four threads, so on a Deck one sleeping fetch is a quarter of the machine's HTTP capacity. Judged by fetch_outstanding and settled_s over a cold cache (CARTO_CACHE_DIR pointed at an empty directory) at a Groningen viewpoint, and by no thread::sleep remaining on an async path in that module."}, {"title": "Report the pipeline census and whether bindless engaged in --dump-state", "note": "systems/dev/pipeline_health.rs already counts ready, building and errored pipelines for the diagnostics report, and quality::Capabilities already records what the adapter answered; neither reaches --dump-state, so no scripted run can say whether geometry was missing because a shader had not compiled yet. Add pipelines_ready, pipelines_building and an errored count to ShotMetrics, plus a flag for whether the adapter offers the texture and buffer binding arrays that StandardMaterial's bindless declaration needs. This is the instrument two later decisions depend on: whether first-run and post-teleport stutter is shader compilation, and whether the 224 per-tile and 160 per-label materials still cost anything on this GPU. Judged by the fields appearing on a settled run and by --expect 'pipelines_errored==0' holding."}, {"title": "Stop submitting shadow casters where no shadow can be resolved", "note": "The cascade configuration in systems/map/sky.rs and its callers. The survey measured 6,998 of 7,978 triangles, 88 per cent, going into shadow cascades at 7,191 km altitude where no shadow is resolvable, and records that nothing gates cascade submission on the tier. At ground level shadow casters were about 79 per cent of submitted triangles, which makes this the largest single geometry setting in the scene on a machine short of both fill rate and memory bandwidth. Gate cascade submission on globe::GlobeState::in_space first, since that case is three passes over invisible geometry. Judged in --dump-state: tris_shadow at an orbit viewpoint falls to zero, while draws and tris_total at a ground viewpoint are unchanged."}, {"title": "Warm the depth-bias pipeline set before the first frame that needs it", "note": "Follows the census ticket, which is what establishes whether this is worth doing. StandardMaterial::depth_bias is part of the specialization key and this tree uses roughly 25 distinct rungs of it, so each is a separate shader compile, and geometry whose pipeline is not yet resident is silently skipped rather than delayed. That costs time-to-a-complete-picture after launch and after every teleport, which on a handheld is most of what a person notices, and --shot hides it by waiting for settle. Land the warm-up alongside the existing PipelineHealthPlugin in the render app. Judged by pipelines_building reading zero at the settle verdict at a viewpoint reached by teleport, with settled_s not rising."}, {"title": "Composite the user interface into the upscaling pass instead of a second full-screen pass", "note": "systems/gui/egui_manager.rs's egui_pass, replacing Bevy's upscaling node with a local one that blits and then hands the still-open RenderPass to egui_wgpu; render_scale::make_the_blit_filter already shows BlitPipeline is reachable from RenderStartup. Today the interface is a full-resolution pass over the swapchain with LoadOp::Load every frame, a whole load and store of the output purely to draw the chrome, which the survey calls the real win here and which is bandwidth on a machine whose GPU shares memory with the CPU. Last on the list because the failure mode is a wgpu validation error, which this tree answers by quitting mid-stream with nothing in the log. Judged by cargo shots producing byte-identical images to the current output, since the harness is deterministic, and by --expect 'draws>0' with an errored-pipeline count of zero."}]

Ticket closed.

No commits were made.


🤖 viberfox issue agent on the VPS — 30 turns · $3.41 · session 174b0941-98df-4d7b-b1ab-946c89de4dc5. Reply with @agent to continue this session.

No code changed and the branch has no commits — this was a planning pass. I read `docs/direction.md`, `CLAUDE.md`, `quality.rs`, `lib.rs`'s task-pool setup, `frame_pacing.rs`, `map_stream.rs` and the performance, streaming and allocation-budget notes. **The main finding:** a Steam Deck is native Linux, so `quality::auto()` hands it `Budget::desktop()` — four shadow cascades to 3 km, 2× multisampling, 512-pixel tile rasters, the sky lighting cube re-baked every frame, and no ceiling at all on either upload rate or how many bytes of streamed map stay resident. `systems::load_control` therefore never runs on it. The rung that already describes the machine in its own documentation is `Budget::balanced` — "an integrated GPU, a thin laptop" — and nothing in the tree can currently reach it on a Deck. Four things I checked that are **already done**, so the queue leaves them alone: the idle frame-rate throttle and the sleep limiter both work on the Linux path; `target-cpu=x86-64-v3` already covers the Deck's Zen 2 cores; coarse basemap tiles fully covered by finer ones are already culled (`map_stream::occluded_by_finer`); and the tile pipeline is already gated on allocation counts. **Left off deliberately:** the globe cloud shader, which the survey says is unmeasurable on this container's software renderer and needs the device; anything about colour, lighting or how the picture reads; and controller input, which is not performance. Every number in these tickets has to be measured headlessly at 1280×800 here. **Nobody in this lane has a Steam Deck**, so confirmation on the actual hardware is yours. ```autopilot-plan [{"title": "Add a handheld render rung and reach it with CARTO_QUALITY=steamdeck", "note": "Lands in crates/cartopolis/src/quality.rs. Add handheld_guarded(Budget::balanced()) in the same shape as android_budget/handset_guarded, with the upload ceiling and resident-bytes target measured rather than guessed, and add the words steamdeck and handheld to parse(). No new QualityPreset variant, so the settings picker and user.json are untouched and the rung is reached by name or by auto-detection. Judged by cargo test -p cartopolis (the existing rung-ladder and tier-vocabulary tests, extended) and by a --shot --size 1280x800 --dump-state pair over Groningen at 150 m and 1700 m: draws, tris_shadow and gpu_resident_kb all below the desktop rung at the same viewpoint, with settled still true. Write the measured table into docs/notes/steam-deck.md the way handset_guarded's table is written."}, {"title": "Answer the handheld rung automatically on a Steam Deck", "note": "quality::auto()'s Linux arm and quality::on_this_machine, which today applies its guards only on Android. quality::init runs before a render adapter exists (the worker pool is sized from its answer), so the adapter cannot be asked; the facts available that early are the machine naming itself: SteamDeck=1 in the environment, and /sys/class/dmi/id/product_name of Jupiter or Galileo with board_vendor Valve. Keep the detection as a pure function over those three strings so it is testable off the device, and route it through on_this_machine so a rung picked by hand in Settings cannot delete a ceiling the machine needs, exactly as the handset does. Judged by unit tests covering both Deck models, a non-Valve Linux machine, and a missing /sys, which must still answer desktop; plus report_budget's chose= field naming the detector in the log."}, {"title": "Size the task pools and the tile-worker pool for four physical cores", "note": "lib.rs, both the TaskPoolThreadAssignmentPolicy block and the platform::workers::start call above it. A Steam Deck is four cores with hyperthreading, so available_parallelism answers 8 and the tile-worker pool takes 7: seven blocking rasterise jobs of roughly 120 ms each, alongside 4 io, 4 async-compute and 3 compute threads, on four physical cores sharing one 15 watt package with the GPU. This is the same arithmetic that silently left Android's parallel system executor one thread wide, and the survey's own lesson is that it was invisible because the widths had only ever been checked on a 22-core machine. Extract the split into a plain function the way frame_pacing::resolve_frame_rate was extracted, and set worker_cap on the handheld rung. Judged by unit tests of that function at 4, 8, 16 and 22 reported threads, and by settled_s and worker_outstanding at a fixed viewpoint not regressing in --dump-state."}, {"title": "Stop three thread::sleep calls from holding io-pool threads in osm_buildings", "note": "crates/cartopolis/src/systems/map/osm_buildings.rs: three paths inside async functions call std::thread::sleep, so each parks an io-pool thread instead of yielding. docs/notes/render-performance-survey.md lists this as open and rejects adding async-io as a direct dependency, but that was written before platform::fetch_gate and the Dispatcher landed on 2026-09-08, which give somewhere to re-rank the work and retry on a later frame instead of waiting on the thread. The io pool floors at four threads, so on a Deck one sleeping fetch is a quarter of the machine's HTTP capacity. Judged by fetch_outstanding and settled_s over a cold cache (CARTO_CACHE_DIR pointed at an empty directory) at a Groningen viewpoint, and by no thread::sleep remaining on an async path in that module."}, {"title": "Report the pipeline census and whether bindless engaged in --dump-state", "note": "systems/dev/pipeline_health.rs already counts ready, building and errored pipelines for the diagnostics report, and quality::Capabilities already records what the adapter answered; neither reaches --dump-state, so no scripted run can say whether geometry was missing because a shader had not compiled yet. Add pipelines_ready, pipelines_building and an errored count to ShotMetrics, plus a flag for whether the adapter offers the texture and buffer binding arrays that StandardMaterial's bindless declaration needs. This is the instrument two later decisions depend on: whether first-run and post-teleport stutter is shader compilation, and whether the 224 per-tile and 160 per-label materials still cost anything on this GPU. Judged by the fields appearing on a settled run and by --expect 'pipelines_errored==0' holding."}, {"title": "Stop submitting shadow casters where no shadow can be resolved", "note": "The cascade configuration in systems/map/sky.rs and its callers. The survey measured 6,998 of 7,978 triangles, 88 per cent, going into shadow cascades at 7,191 km altitude where no shadow is resolvable, and records that nothing gates cascade submission on the tier. At ground level shadow casters were about 79 per cent of submitted triangles, which makes this the largest single geometry setting in the scene on a machine short of both fill rate and memory bandwidth. Gate cascade submission on globe::GlobeState::in_space first, since that case is three passes over invisible geometry. Judged in --dump-state: tris_shadow at an orbit viewpoint falls to zero, while draws and tris_total at a ground viewpoint are unchanged."}, {"title": "Warm the depth-bias pipeline set before the first frame that needs it", "note": "Follows the census ticket, which is what establishes whether this is worth doing. StandardMaterial::depth_bias is part of the specialization key and this tree uses roughly 25 distinct rungs of it, so each is a separate shader compile, and geometry whose pipeline is not yet resident is silently skipped rather than delayed. That costs time-to-a-complete-picture after launch and after every teleport, which on a handheld is most of what a person notices, and --shot hides it by waiting for settle. Land the warm-up alongside the existing PipelineHealthPlugin in the render app. Judged by pipelines_building reading zero at the settle verdict at a viewpoint reached by teleport, with settled_s not rising."}, {"title": "Composite the user interface into the upscaling pass instead of a second full-screen pass", "note": "systems/gui/egui_manager.rs's egui_pass, replacing Bevy's upscaling node with a local one that blits and then hands the still-open RenderPass to egui_wgpu; render_scale::make_the_blit_filter already shows BlitPipeline is reachable from RenderStartup. Today the interface is a full-resolution pass over the swapchain with LoadOp::Load every frame, a whole load and store of the output purely to draw the chrome, which the survey calls the real win here and which is bandwidth on a machine whose GPU shares memory with the CPU. Last on the list because the failure mode is a wgpu validation error, which this tree answers by quitting mid-stream with nothing in the log. Judged by cargo shots producing byte-identical images to the current output, since the harness is deterministic, and by --expect 'draws>0' with an errored-pipeline count of zero."}] ``` Ticket closed. _No commits were made._ --- <sub>🤖 viberfox issue agent on the VPS — 30 turns · $3.41 · session `174b0941-98df-4d7b-b1ab-946c89de4dc5`. Reply with **@agent** to continue this session.</sub>
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
jeroen/cartopolis#292
No description provided.