diff options
| author | Christophe Besson <cbesson@gmail.com> | 2026-10-02 11:54:56 +0200 |
|---|---|---|
| committer | Christophe Besson <cbesson@gmail.com> | 2026-10-02 11:54:56 +0200 |
| commit | d7b7f1049d95e45e6316ac419cb434088c15bd5e (patch) | |
| tree | 77da46e2fe3cb70af5fb2eebcbb1d69dad410b88 /docs/MESHBAY_DESIGN.md | |
| parent | 56a8cf9167e8c7b0f2df15afed88031608adf782 (diff) | |
| parent | 754387590fa1754436b4648f969915888c6f6c9e (diff) | |
| download | meshbay-d7b7f1049d95e45e6316ac419cb434088c15bd5e.tar.gz | |
Diffstat (limited to 'docs/MESHBAY_DESIGN.md')
| -rw-r--r-- | docs/MESHBAY_DESIGN.md | 69 |
1 files changed, 39 insertions, 30 deletions
diff --git a/docs/MESHBAY_DESIGN.md b/docs/MESHBAY_DESIGN.md index eec354a..189dbc5 100644 --- a/docs/MESHBAY_DESIGN.md +++ b/docs/MESHBAY_DESIGN.md @@ -16,7 +16,7 @@ > them — it names the invariant that holds today, not the incident that produced > it. §13 is the register of those labels. > -> Wire versions at the time of writing: **MNP 5.0** (oldest peer accepted 4.0), +> Wire versions at the time of writing: **MNP 6.0** (oldest peer accepted 4.0), > **MHP 0.1**, packages **0.17.0**. The normative source for the wire format is > `MESHBAY_NODE_PROTOCOL.md`; this document states the design the protocol > serves, not its byte layout. @@ -99,7 +99,7 @@ opens them. │ └─────────┘ MHP 0.1 │ signalling (SDP/ICE, <1 KB), presence, revocation push MHP │ - ┌────┴────┐ MNP 5.0 ┌──────────┐ + ┌────┴────┐ MNP 6.0 ┌──────────┐ │ node │◄──────── WebRTC DataChannel / QUIC ──────────►│ client │ └─────────┘ index, file chunks, streams, chat, admin └──────────┘ holds the files browser SPA or desktop @@ -786,10 +786,10 @@ Three properties are why this shape: The node produces every copy of the key itself, from its own CSPRNG. **Nothing arriving over MNP can activate a group key** (**C5b**). Read that precisely: it -targets *key material arriving from outside*, not the instruction. An -operator-signed `gek_rotate` where the node generates the key is a different shape -and is allowed. The initial `gek-init` stays local, because with no key there is -no completed session to carry a signed op. +targets *key material arriving from outside*, not the instruction: a key the node +generates itself on an operator's instruction is a different shape. Initialising +and rotating the group key are local (loopback API, CLI); the signed `chat_epoch` +is the MNP instance of that shape. ### 4.3 On-the-fly encryption @@ -898,7 +898,7 @@ signature refuses it. **Epochs.** A new epoch is opened when, and only when, the set of devices that may read *future* messages shrinks: `member revoke`, `member unpin`, `revoke_device`, -`gek_rotate`, or an explicit `chat rotate`. Epoch 1 is opened at group load — a +or an explicit `chat rotate`. Epoch 1 is opened at group load — a group with no epoch is a group nobody can speak in. **Old epochs are kept and still delivered.** That is what keeps history readable @@ -1233,9 +1233,9 @@ from anything in the response. | `dir_delete` | the operator alone, and only on an empty directory | | `invite_create` | the operator (or a delegate, when delegation ships) | | `invite_link_create`, `invite_cancel` | the operator | -| `gek_rotate` | operator-signed; the node generates the key itself | -| initial `gek-init` | **local admin API or CLI only** | -| root add/remove/update/eject/plug, `apps_enabled`, app directories, transfer limits | operator-signed | +| group key init and rotation | **local admin API or CLI only**; the node generates the key itself | +| root add, root `writable`/`removable`, hosting a group, transfer limits | **local admin API or CLI only** (MNP 6.0) | +| root remove/eject/plug, `apps_enabled`, app directories | operator-signed | | ~~`gek_bundle_store`~~ | **the message does not exist.** No member ever hands the node key material | `gek_bundle_store` was deleted rather than gated. The operator's X25519 public key @@ -1451,8 +1451,8 @@ checks the version its peer declared and **branches on none of it**. **The floor is not necessarily the current version, and what is added above it is why.** It is `MNP_MIN_SUPPORTED` in `handshake.py`, and it is the last MAJOR that had to refuse at the handshake: 3.1–3.4 were added above the 3.0 floor without moving it, -and 5.0, a MAJOR confined to four signed operations that a peer across the break -refuses to sign, sits above the 4.0 floor. So a +and 5.0 and 6.0 — MAJORs confined to a few signed operations, four whose subjects +changed and then three removed — sit above the 4.0 floor. So a peer can be reachable and still not do something the current version can, and the client has to cope with that — **by reading the peer's own answer, never by comparing version numbers**. @@ -1541,7 +1541,16 @@ Five consequences, none optional: - `writable = true` means any group member may upload there. Several roots may be writable and none need be — a fully read-only group is valid. -The operator toggles this with a signed op. + +**What widens the sharing is decided on the node's own machine.** Adding a root, +hosting a group over a directory, and switching `writable` or `removable` go through +the loopback API (the desktop application) or the CLI — never over MNP, since 6.0. +A signed op proves that the operator's key signed, not that they meant it: in a +browser that key is driven by code the hub serves (T3), and in the desktop +application by a renderer that parses content from nodes. Either could otherwise +have shared any folder on the machine, writable, from anywhere. The operator still +sees every root and its flags from any browser; removing, ejecting and plugging stay +signed ops, because they narrow what is shared or restore what already was. > **There is one answer to "may this member write", and it is the root.** A single > flag over the group cannot express "this library is published read-only and that @@ -1890,26 +1899,24 @@ issues invitations and reads the audit log. There is no server-rendered dashboar the desktop client's Node page and the CLI are the two consumers, and each operation endpoint is one `_op(...)` line onto `ops` (§5.4). -**Over MNP, the node's own controls need a proved operator device.** The -node-wide surface — `node_status`, which lists every group on the machine with -each root's absolute path, plus `node_settings_set`, `roster_read`, -`denylist_read`, `denylist_clear` and `node_reload` — is reachable when two -things hold: the account is the one the node belongs to, *and* the device on the -connection has proved (`device_hello`, §3.3) a key the roster holds as an -operator. The first alone is a claim in a token the hub issued, and **NS4** does -not allow it to be authority: a hub that can name the operator is a hub that can -be one. The second is what it cannot forge, since it holds no user keys and -cannot countersign a device — the same property device linking rests on. A -browser that has never been paired therefore reads nothing here, exactly as it -can already sign nothing (§5.4). +**The node's own controls are not on MNP.** Its status — which lists every group +on the machine with each root's absolute path — settings, roster, denylist and +reload were MNP messages gated on a proved operator device (`device_hello`, §3.3), +because the account id in a token is the hub's to choose (**NS4**). No client ever +sent them, and MNP 6.0 removed them with the four signed ops in the same position +(`gek_rotate`, `member_unpin`, `transfer_limits`, `group_detach`): the desktop +client's Node page and the CLI do this work over loopback. A door nobody calls is +an untested way in, and one that does not exist needs no gate. **The accepted cost, recorded as a choice:** on a headless server the only admin path is the CLI. The CLI covers every operation, so this is acceptable — but it is a real capability reduction, not an oversight. -> **MNP is the path that must exist; loopback is the fallback.** The operator of a -> node is not necessarily sitting at it. Any operator-facing control needs its MNP -> route first, or it renders for nobody on the web. +> **What the operator must see needs an MNP route; what widens the node does not +> get one.** The operator of a node is not necessarily sitting at it, so a view that +> only loopback can fill renders for nobody on the web — which is why the roots +> table rides in the index. But sharing a folder, opening it to writes and the +> node's own controls are decided at the node (§6.2, MNP 6.0). ### 6.8 Node settings @@ -3612,7 +3619,7 @@ had already been asked. | **AV15** | **A hash is checked for shape before it is a key lookup**, on every blocklist endpoint, the administrator's included | | **AV20** | **Chat is bounded in size and in rate, like every other member-supplied write** (§6.6). A message is a row on the operator's disk that nothing expires, a relayed copy for every connected member and a notification for every member of the group; the only ceiling was the frame size. Uploads had carried four protections and a cap since C5a because somebody asked what one member costs the others on that path, and nobody had asked it on this one | | **AV21** | **A lease is what the node granted, not what the client called it** (§5.5). `tr` was read as a boolean, so any non-empty string skipped the leaseless ceiling and every cap behind it, and a queued transfer was held back only by the honesty of the client waiting in the queue | -| **AV22** | **The node's own controls take no authority from a hub token** (§6.7). `node_status`, `node_settings_set`, `roster_read`, `denylist_read`, `denylist_clear` and `node_reload` were gated on the account id in the JWT, which is the hub's to choose — NS4 and M3 with the check written the other way round. The gate is a proved operator device, which a hub holding no user keys cannot produce | +| **AV22** | **The node's own controls take no authority from a hub token** (§6.7). `node_status`, `node_settings_set`, `roster_read`, `denylist_read`, `denylist_clear` and `node_reload` were gated on the account id in the JWT, which is the hub's to choose — NS4 and M3 with the check written the other way round. The gate became a proved operator device, which a hub holding no user keys cannot produce; since MNP 6.0 the messages are gone and these controls are loopback and CLI only | | **AV23** | **An upload's owner is recorded when the upload ends and applied when the entry is created**, which are different moments (§5.4). Written against the index at the end of the upload it matched nothing, every time, and left every uploaded file owned by nobody — so no member could delete what they had sent | | **AV24** | **A node registered for no group is refused signaling, not exempted from it** (§7.2). The membership check was written as "if the node claims any group", so it skipped itself — membership, group status and the public-group gate together — for the node AV1 made commonplace: the unconfigured one, which is also the one least able to absorb the work | | **AV25** | **Which nodes host a group is answered to its members** (§7.3). Only the public case checked, so a private group told any authenticated account that knew its id which machines hosted it — and an ex-member knows that id for ever | @@ -3798,10 +3805,12 @@ process runs it — `systemctl --user` on Linux, Task Scheduler on Windows. | **The reconnect backoff only wakes on `visibilitychange`** | So a tab that stays visible through an outage — which is what a screen wake lock guarantees while a film is playing — waits out the full backoff, up to 30 s, after the network is already back. Nothing listens for `online` | | **Per-device revocation has no CLI** | A device is revoked over MNP (`roster.revoke_device`), from a device the node has already pinned. On a headless node the operator's only lever is `member unpin`, which removes **every** device of that account — so the per-device control the roster is built around is reachable from an interface and from nowhere else. §6.7 listed a `meshbay-node member device list\|revoke` verb that was never written, and that listing is how this was found: `USERGUIDE.md` was the first document written by reading the CLI rather than this specification, and the verb it copied out did not run | | **Migrations run on SQLite only** | The chain reaches head and agrees with the models there (§12), which is not where it ships. **The exposure is one revision deep, not the whole chain**: every revision behind the first packaged release was development that no installation ever ran, so nothing replays them on PostgreSQL. What is unguarded is the *next* migration — a default, an index type or a constraint PostgreSQL refuses reaches a deploy without the suite saying so | -| **The loopback path removes access without writing an audit entry** | `ops.revoke_member` and `ops.unpin_member` log to the daemon's log and nothing to `audit.db`; the MNP admin handlers doing the same work audit `member_revoke` and `member_unpin`. So a removal made from the node page or the CLI — the two doors an operator sitting at their own machine actually uses — leaves the journal showing an admission and then, whenever that person next connects, an `auth_failed` ("not admitted by the roster") with nothing in between to explain it. §5.4's signed transcript is not what is missing: a loopback caller is authorized by being on localhost with the run token and signs nothing, so the gap is the record, not the authority. Found by reading a node's audit log for a refusal whose cause was six hours earlier and unrecorded | +| **The loopback path removes access without writing an audit entry** | `ops.revoke_member` and `ops.unpin_member` log to the daemon's log and nothing to `audit.db`; the MNP `member_revoke` handler doing the same work audits it (and `member_unpin` did, until it left MNP in 6.0). So a removal made from the node page or the CLI — the two doors an operator sitting at their own machine actually uses — leaves the journal showing an admission and then, whenever that person next connects, an `auth_failed` ("not admitted by the roster") with nothing in between to explain it. §5.4's signed transcript is not what is missing: a loopback caller is authorized by being on localhost with the run token and signs nothing, so the gap is the record, not the authority. Found by reading a node's audit log for a refusal whose cause was six hours earlier and unrecorded | | **The Create group wizard calls two hooks after an early return** | `CreateGroupWizard` (`create-group-page.js`) returns during node detection, before its `useRef`/`useEffect` for provisioning, so the hook count changes between renders. Preact tolerates a list that grows, and nothing is known to break; `test_hook_ordering.py` checks declaration order, not this. Found while tracing the frozen-fields report, which had another cause (`ask.js`) | | **A node key is read from the terminal or the desktop client, never a browser** | **Accepted.** `meshbay-node status` on the node's own machine and Node → Overview in the desktop client are the two places the key can be read; the Node page is Electron-only, because `platform.node` resolves to "not available" without the bridge, and no hub route exposes the key. The create-group wizard links it automatically over that same bridge, so the manual paste in **Profile → Link Node** exists for the operator who runs the node from a terminal and the hub from a browser — who has a terminal by definition. Anyone linking a node is already at a shell prompt, so a browser-reachable copy would buy nothing and widen what the hub knows about the node | | **Listing a group's folders walks every root on the event loop** | `index_sync_message` (`transport/wire.py`) builds its `dirs` field with `list_dirs`, an `rglob("*")` over every root, and nothing sends it off the loop: the WebRTC `index_sync` handler, the daemon's index push and QUIC all call it inline. So each index request from any member is a directory walk of the whole library that every other peer on the node waits behind. `test_disk_io_off_loop.py` never saw it, because it reads the transport's own modules and the walk is one call away in `wire.py`. Found by widening what that test reads, not by a symptom | +| **A loopback eject or plug reaches open pages late** | `ops.eject_root` and `ops.plug_root` flip the live set and tell nobody; over MNP the broadcast `root_eject_ack` / `root_plug_ack` is what moves every open table. So an eject made from the desktop application or the CLI shows on members' pages only with the next index push — for a plug, the end of its rescan; for an eject, whatever changes next. `ops.update_root` had the same silence and now calls `DirectoryIndexer.publish_roots`; the same call belongs in these two. Found while moving the `writable`/`removable` switches to the loopback door (MNP 6.0) | +| **A group key rotated from the Node page leaves the chat key where it was** | `ops.set_gek(rotate=True)` replaces the group key and opens no chat epoch; the MNP `gek_rotate` handler opened one itself (`_new_chat_epoch`), and it was the only door that did — but no client ever sent it, and it is gone since 6.0. The removals that matter (revoke, unpin, device revoke) open an epoch in `ops` for every door, so what is missing is the follow-through for an operator who rotates by hand: §4.5's "rotate after a removal" means it for chat too. The fix is the `_after_removal` shape — `open_chat_epoch` inside `ops.set_gek` when `rotated`. Found while removing the MNP message | | **The transcoded-seek test passes without transcoding** | `test_a_transcoded_video_keeps_accurate_seeking` (`test_stream_seek_audio_alignment.py`) forces the re-encode branch by swapping the module's `BROWSER_INCOMPATIBLE_VIDEO_CODECS`, then checks only that the result has no audio gap. The copy path also leaves no gap on that clip, so pointing the swap at a module the streaming code does not read still passes: the test cannot tell that the branch it is named after never ran. It should assert the re-encode happened (the `re-encoding` log line, or the encoder in the ffmpeg argv). Found by breaking the swap on purpose while moving the streaming code | --- |