gap: a cell's ground waits on api.pdok.nl before it is tessellated #275

Merged
viberfox-agent merged 1 commit from fix/267-bound-optional-layer-fetches into main 2026-09-02 08:28:44 +00:00
Collaborator

Closes #267

Merged by the autopilot (docs/direction.md) after every Actions job passed on the branch head.

Closes #267 Merged by the autopilot (docs/direction.md) after every Actions job passed on the branch head.
fix(map): bound the two optional PDOK fetches a cell's ground waits on
All checks were successful
CI / test cartopolis (pull_request) Successful in 11m9s
CI / wasm & android targets (pull_request) Has been skipped
16091087d6
`spawn_geometry_task` awaits the navigation marks and the level crossings
between the cell's tile body and `workers::submit`, so a cell's ground fills,
roads, water, bridges, trees, labels and building footprints were all held
behind `api.pdok.nl` — a host that owns none of them. Both built a plain
`Request::get`, i.e. the inherited 20 s `DEFAULT_TIMEOUT`, and the native fetch
is a blocking `ureq` call that holds one of at most eight `IoTaskPool` threads
for the whole of it. With eight surface cells in flight, a PDOK incident stops
the basemap and the building streamers too. Healthy the layers are cheap: 18
requests over Groningen took 1.70 s, median 82 ms, worst 186 ms. The unhealthy
end is what had no bound.

Two bounds, both in `platform::http` so a third such layer copies nothing.

`OPTIONAL_LAYER_TIMEOUT` (3 s, ~16x the measured worst case) is carried by
every request in `nav_marks` and `crossings`: a slow-but-working service still
delivers, a dead one costs three seconds instead of twenty.

`http::Breaker` remembers the failure, which nothing did before. Neither layer
caches a failed answer, deliberately — a poisoned key is never re-fetched — and
`map_geometry` re-runs the whole task on a re-anchor, a terrain-field revision,
a coverage revision or any detail-layer toggle, so every retear re-paid the
full wait for every cell in the ring. Any fetch or parse failure now trips that
service's breaker and `should_fetch` answers `false` for 60 s, matching
`tall_structures::RETRY_AFTER_SECS`. One instance per service, not per host:
RWS publishes the marks and ProRail the crossings, and they fail apart. Its
clock is `web_time` (`std::time::Instant::now()` panics on wasm, which is the
one platform where the breaker is the only bound, `ehttp` ignoring the request
timeout), and the decision is a pure predicate over an injected instant so both
gates are asserted with no runtime, no network and no sleeping out a cooldown.

The gates are split rather than widened: `cell_may_hold_marks` /
`cell_may_hold_crossings` are the cell's own two, checked before the cache;
`should_fetch` adds the breaker and is checked after it. A cell already on disk
must cost the service nothing, so an outage cannot blank cells that were
complete before it started, and a foreign or dry cell still costs neither a
request nor a storage lookup.

The two collections and the two layers are awaited through
`futures_lite::future::zip`. That is real overlap on wasm only: natively
`http::fetch` has no await point in it, so the requests still land in sequence
there and the timeout plus the breaker are what cap the wait. Written down as
such rather than claimed as concurrency.

Refs #267
viberfox-agent deleted branch fix/267-bound-optional-layer-fetches 2026-09-02 08:28:45 +00:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
jeroen/cartopolis!275
No description provided.