Proximity voice over the existing wire protocol (sim side) #32

Closed
opened 2026-08-10 11:55:22 +00:00 by jeroen · 0 comments
Owner

Corrected 2026-08-10. The first version of this spec cited crates/vibe_core
and crates/vibe_sim, which have not existed since the crate rename. The work
was rebuilt on origin/main as viberfox_core / viberfox_simulator; paths
below are the real ones. Branch: feat/32-proximity-voice-sim.

Problem

The sim has text chat but no voice. crates/simulator/src/chat.rs (85 lines) is
already the routing fabric spatial voice needs — it deals in pre-encoded frames
(mpsc::UnboundedSender<Vec<u8>>, chat.rs:18), and already does the three
deliveries voice wants: deliver_all, deliver_room, deliver_to
(chat.rs:59-81), with proximity radii already chosen (SAY_RADIUS_M = 50.0,
SHOUT_RADIUS_M = 300.0, chat.rs:14-15).

Decisively, chat.rs:3-4 records that proximity (Local) already encodes per
recipient
because each frame carries its own distance_m. Per-recipient
framing is exactly what distance attenuation needs — the design decision is made,
implemented and documented, for text.

ADR-028:276-289 leans toward str0m and an embedded SFU. That is superseded: an
SFU reimplements ChatHub's job while knowing nothing about avatar positions,
which then have to be fed to it.

Approach

Server and protocol only — the half that needs no audio device and no browser
transport.

  • crates/core/src/protocol.rs — VoiceFrame (client→server) and VoiceData
    (server→client) at kinds 20/21, reusing ChatScope unchanged. One appended
    ProtocolRevision (13); the version is derived from it.
  • crates/simulator/src/chat.rs — no generalising was needed: ChatHub already
    deals in encoded Vec<u8>, so its payload was always opaque. Adds
    deliver_room_except and extracts within_radius.
  • crates/simulator/src/net.rs — handle the new kinds; voice rides the existing
    Link and its existing auth. No second connection.
  • Rate limits sized for media, not typing.

Acceptance criteria

  • PROTOCOL_HISTORY gains exactly one row; the gapless-run test still passes.
  • A voice frame sent Local reaches only avatars within SAY_RADIUS_M, and
    shout extends that to SHOUT_RADIUS_M.
  • Each recipient's frame carries its own distance_m, as text already does.
  • Voice payloads are never written to SQLite scrollback.
  • An unauthenticated connection cannot send or receive voice.
  • Text chat behaviour is unchanged — crates/simulator/tests/chat.rs passes
    untouched.

Added while building, each with a test:

  • Voice never loops back to the speaker (an echo, where text's echo is
    confirmation).
  • No global voice tier — unbounded fan-out is the loudest abuse surface.
    One line to enable if that call is revisited.
  • VoiceFrame carries seq, because the intended transport is unreliable,
    unordered datagrams where arrival order is not send order.
  • An oversize payload is refused without killing the link.

Verification

cargo test -p viberfox_core (20 passed) and cargo test -p viberfox_simulator
(24 passed across lib + 3 integration targets), plus cargo check -p viberfox.

The radius is covered by a unit test in chat::tests rather than through the
socket: both test clients authenticate as the same account and spawn co-located,
and SimHandle exposes no way to move them apart.

Nothing audible is verified. The agent container has no audio device and no
microphone. pcm.null exists in libasound but does not clock, so it cannot
exercise callback pacing; a properly-clocked fake needs either a PulseAudio
null-sink in the agent image or snd-aloop on the host.

Out of scope

  • Audio I/O (platform::audio), Opus, and acoustic echo cancellation.
    AEC decides whether DIY voice sounds professional; getUserMedia supplies it on
    web, native is ours. Separate ticket.
  • The datagram browser transport (WebTransport) — the real prerequisite for
    voice actually flowing on web, valuable alone, and its own ticket.
  • Video and file sharing.
  • Any SFU, ICE, DTLS/SRTP or TURN.
  • ChatHub's Mutex<HashMap> + per-recipient frame.to_vec() is left as-is.
    Free at chat rates; wants measuring at 50 packets/sec × N listeners. That
    measurement belongs in a note, not this ticket.
> **Corrected 2026-08-10.** The first version of this spec cited `crates/vibe_core` > and `crates/vibe_sim`, which have not existed since the crate rename. The work > was rebuilt on `origin/main` as `viberfox_core` / `viberfox_simulator`; paths > below are the real ones. Branch: `feat/32-proximity-voice-sim`. ## Problem The sim has text chat but no voice. `crates/simulator/src/chat.rs` (85 lines) is already the routing fabric spatial voice needs — it deals in pre-encoded frames (`mpsc::UnboundedSender<Vec<u8>>`, `chat.rs:18`), and already does the three deliveries voice wants: `deliver_all`, `deliver_room`, `deliver_to` (`chat.rs:59-81`), with proximity radii already chosen (`SAY_RADIUS_M = 50.0`, `SHOUT_RADIUS_M = 300.0`, `chat.rs:14-15`). Decisively, `chat.rs:3-4` records that proximity (`Local`) already encodes **per recipient** because each frame carries its own `distance_m`. Per-recipient framing is exactly what distance attenuation needs — the design decision is made, implemented and documented, for text. ADR-028:276-289 leans toward str0m and an embedded SFU. That is superseded: an SFU reimplements `ChatHub`'s job while knowing nothing about avatar positions, which then have to be fed to it. ## Approach Server and protocol only — the half that needs no audio device and no browser transport. * `crates/core/src/protocol.rs` — `VoiceFrame` (client→server) and `VoiceData` (server→client) at kinds 20/21, reusing `ChatScope` unchanged. One appended `ProtocolRevision` (13); the version is derived from it. * `crates/simulator/src/chat.rs` — no generalising was needed: `ChatHub` already deals in encoded `Vec<u8>`, so its payload was always opaque. Adds `deliver_room_except` and extracts `within_radius`. * `crates/simulator/src/net.rs` — handle the new kinds; voice rides the existing `Link` and its existing auth. No second connection. * Rate limits sized for media, not typing. ## Acceptance criteria - [x] `PROTOCOL_HISTORY` gains exactly one row; the gapless-run test still passes. - [x] A voice frame sent `Local` reaches only avatars within `SAY_RADIUS_M`, and `shout` extends that to `SHOUT_RADIUS_M`. - [x] Each recipient's frame carries its own `distance_m`, as text already does. - [x] Voice payloads are never written to SQLite scrollback. - [x] An unauthenticated connection cannot send or receive voice. - [x] Text chat behaviour is unchanged — `crates/simulator/tests/chat.rs` passes untouched. Added while building, each with a test: - [x] Voice never loops back to the speaker (an echo, where text's echo is confirmation). - [x] No global voice tier — unbounded fan-out is the loudest abuse surface. One line to enable if that call is revisited. - [x] `VoiceFrame` carries `seq`, because the intended transport is unreliable, unordered datagrams where arrival order is not send order. - [x] An oversize payload is refused without killing the link. ## Verification `cargo test -p viberfox_core` (20 passed) and `cargo test -p viberfox_simulator` (24 passed across lib + 3 integration targets), plus `cargo check -p viberfox`. The radius is covered by a unit test in `chat::tests` rather than through the socket: both test clients authenticate as the same account and spawn co-located, and `SimHandle` exposes no way to move them apart. **Nothing audible is verified.** The agent container has no audio device and no microphone. `pcm.null` exists in libasound but does not clock, so it cannot exercise callback pacing; a properly-clocked fake needs either a PulseAudio null-sink in the agent image or `snd-aloop` on the host. ## Out of scope * **Audio I/O** (`platform::audio`), **Opus**, and **acoustic echo cancellation**. AEC decides whether DIY voice sounds professional; `getUserMedia` supplies it on web, native is ours. Separate ticket. * **The datagram browser transport** (WebTransport) — the real prerequisite for voice actually flowing on web, valuable alone, and its own ticket. * Video and file sharing. * Any SFU, ICE, DTLS/SRTP or TURN. * `ChatHub`'s `Mutex<HashMap>` + per-recipient `frame.to_vec()` is left as-is. Free at chat rates; wants measuring at 50 packets/sec × N listeners. That measurement belongs in a note, not this ticket.
Sign in to join this conversation.
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
jeroen/cartopolis#32
No description provided.