summaryrefslogtreecommitdiffstats
path: root/docs/MESHBAY_DESIGN.md
diff options
context:
space:
mode:
Diffstat (limited to 'docs/MESHBAY_DESIGN.md')
-rw-r--r--docs/MESHBAY_DESIGN.md49
1 files changed, 24 insertions, 25 deletions
diff --git a/docs/MESHBAY_DESIGN.md b/docs/MESHBAY_DESIGN.md
index 49bd99b..189dbc5 100644
--- a/docs/MESHBAY_DESIGN.md
+++ b/docs/MESHBAY_DESIGN.md
@@ -786,10 +786,10 @@ Three properties are why this shape:
The node produces every copy of the key itself, from its own CSPRNG. **Nothing
arriving over MNP can activate a group key** (**C5b**). Read that precisely: it
-targets *key material arriving from outside*, not the instruction. An
-operator-signed `gek_rotate` where the node generates the key is a different shape
-and is allowed. The initial `gek-init` stays local, because with no key there is
-no completed session to carry a signed op.
+targets *key material arriving from outside*, not the instruction: a key the node
+generates itself on an operator's instruction is a different shape. Initialising
+and rotating the group key are local (loopback API, CLI); the signed `chat_epoch`
+is the MNP instance of that shape.
### 4.3 On-the-fly encryption
@@ -898,7 +898,7 @@ signature refuses it.
**Epochs.** A new epoch is opened when, and only when, the set of devices that may
read *future* messages shrinks: `member revoke`, `member unpin`, `revoke_device`,
-`gek_rotate`, or an explicit `chat rotate`. Epoch 1 is opened at group load — a
+or an explicit `chat rotate`. Epoch 1 is opened at group load — a
group with no epoch is a group nobody can speak in.
**Old epochs are kept and still delivered.** That is what keeps history readable
@@ -1233,9 +1233,9 @@ from anything in the response.
| `dir_delete` | the operator alone, and only on an empty directory |
| `invite_create` | the operator (or a delegate, when delegation ships) |
| `invite_link_create`, `invite_cancel` | the operator |
-| `gek_rotate` | operator-signed; the node generates the key itself |
-| initial `gek-init` | **local admin API or CLI only** |
-| root add/remove/update/eject/plug, `apps_enabled`, app directories, transfer limits | operator-signed |
+| group key init and rotation | **local admin API or CLI only**; the node generates the key itself |
+| root add, root `writable`/`removable`, hosting a group, transfer limits | **local admin API or CLI only** (MNP 6.0) |
+| root remove/eject/plug, `apps_enabled`, app directories | operator-signed |
| ~~`gek_bundle_store`~~ | **the message does not exist.** No member ever hands the node key material |
`gek_bundle_store` was deleted rather than gated. The operator's X25519 public key
@@ -1899,26 +1899,24 @@ issues invitations and reads the audit log. There is no server-rendered dashboar
the desktop client's Node page and the CLI are the two consumers, and each
operation endpoint is one `_op(...)` line onto `ops` (§5.4).
-**Over MNP, the node's own controls need a proved operator device.** The
-node-wide surface — `node_status`, which lists every group on the machine with
-each root's absolute path, plus `node_settings_set`, `roster_read`,
-`denylist_read`, `denylist_clear` and `node_reload` — is reachable when two
-things hold: the account is the one the node belongs to, *and* the device on the
-connection has proved (`device_hello`, §3.3) a key the roster holds as an
-operator. The first alone is a claim in a token the hub issued, and **NS4** does
-not allow it to be authority: a hub that can name the operator is a hub that can
-be one. The second is what it cannot forge, since it holds no user keys and
-cannot countersign a device — the same property device linking rests on. A
-browser that has never been paired therefore reads nothing here, exactly as it
-can already sign nothing (§5.4).
+**The node's own controls are not on MNP.** Its status — which lists every group
+on the machine with each root's absolute path — settings, roster, denylist and
+reload were MNP messages gated on a proved operator device (`device_hello`, §3.3),
+because the account id in a token is the hub's to choose (**NS4**). No client ever
+sent them, and MNP 6.0 removed them with the four signed ops in the same position
+(`gek_rotate`, `member_unpin`, `transfer_limits`, `group_detach`): the desktop
+client's Node page and the CLI do this work over loopback. A door nobody calls is
+an untested way in, and one that does not exist needs no gate.
**The accepted cost, recorded as a choice:** on a headless server the only admin
path is the CLI. The CLI covers every operation, so this is acceptable — but it is
a real capability reduction, not an oversight.
-> **MNP is the path that must exist; loopback is the fallback.** The operator of a
-> node is not necessarily sitting at it. Any operator-facing control needs its MNP
-> route first, or it renders for nobody on the web.
+> **What the operator must see needs an MNP route; what widens the node does not
+> get one.** The operator of a node is not necessarily sitting at it, so a view that
+> only loopback can fill renders for nobody on the web — which is why the roots
+> table rides in the index. But sharing a folder, opening it to writes and the
+> node's own controls are decided at the node (§6.2, MNP 6.0).
### 6.8 Node settings
@@ -3621,7 +3619,7 @@ had already been asked.
| **AV15** | **A hash is checked for shape before it is a key lookup**, on every blocklist endpoint, the administrator's included |
| **AV20** | **Chat is bounded in size and in rate, like every other member-supplied write** (§6.6). A message is a row on the operator's disk that nothing expires, a relayed copy for every connected member and a notification for every member of the group; the only ceiling was the frame size. Uploads had carried four protections and a cap since C5a because somebody asked what one member costs the others on that path, and nobody had asked it on this one |
| **AV21** | **A lease is what the node granted, not what the client called it** (§5.5). `tr` was read as a boolean, so any non-empty string skipped the leaseless ceiling and every cap behind it, and a queued transfer was held back only by the honesty of the client waiting in the queue |
-| **AV22** | **The node's own controls take no authority from a hub token** (§6.7). `node_status`, `node_settings_set`, `roster_read`, `denylist_read`, `denylist_clear` and `node_reload` were gated on the account id in the JWT, which is the hub's to choose — NS4 and M3 with the check written the other way round. The gate is a proved operator device, which a hub holding no user keys cannot produce |
+| **AV22** | **The node's own controls take no authority from a hub token** (§6.7). `node_status`, `node_settings_set`, `roster_read`, `denylist_read`, `denylist_clear` and `node_reload` were gated on the account id in the JWT, which is the hub's to choose — NS4 and M3 with the check written the other way round. The gate became a proved operator device, which a hub holding no user keys cannot produce; since MNP 6.0 the messages are gone and these controls are loopback and CLI only |
| **AV23** | **An upload's owner is recorded when the upload ends and applied when the entry is created**, which are different moments (§5.4). Written against the index at the end of the upload it matched nothing, every time, and left every uploaded file owned by nobody — so no member could delete what they had sent |
| **AV24** | **A node registered for no group is refused signaling, not exempted from it** (§7.2). The membership check was written as "if the node claims any group", so it skipped itself — membership, group status and the public-group gate together — for the node AV1 made commonplace: the unconfigured one, which is also the one least able to absorb the work |
| **AV25** | **Which nodes host a group is answered to its members** (§7.3). Only the public case checked, so a private group told any authenticated account that knew its id which machines hosted it — and an ex-member knows that id for ever |
@@ -3807,11 +3805,12 @@ process runs it — `systemctl --user` on Linux, Task Scheduler on Windows.
| **The reconnect backoff only wakes on `visibilitychange`** | So a tab that stays visible through an outage — which is what a screen wake lock guarantees while a film is playing — waits out the full backoff, up to 30 s, after the network is already back. Nothing listens for `online` |
| **Per-device revocation has no CLI** | A device is revoked over MNP (`roster.revoke_device`), from a device the node has already pinned. On a headless node the operator's only lever is `member unpin`, which removes **every** device of that account — so the per-device control the roster is built around is reachable from an interface and from nowhere else. §6.7 listed a `meshbay-node member device list\|revoke` verb that was never written, and that listing is how this was found: `USERGUIDE.md` was the first document written by reading the CLI rather than this specification, and the verb it copied out did not run |
| **Migrations run on SQLite only** | The chain reaches head and agrees with the models there (§12), which is not where it ships. **The exposure is one revision deep, not the whole chain**: every revision behind the first packaged release was development that no installation ever ran, so nothing replays them on PostgreSQL. What is unguarded is the *next* migration — a default, an index type or a constraint PostgreSQL refuses reaches a deploy without the suite saying so |
-| **The loopback path removes access without writing an audit entry** | `ops.revoke_member` and `ops.unpin_member` log to the daemon's log and nothing to `audit.db`; the MNP admin handlers doing the same work audit `member_revoke` and `member_unpin`. So a removal made from the node page or the CLI — the two doors an operator sitting at their own machine actually uses — leaves the journal showing an admission and then, whenever that person next connects, an `auth_failed` ("not admitted by the roster") with nothing in between to explain it. §5.4's signed transcript is not what is missing: a loopback caller is authorized by being on localhost with the run token and signs nothing, so the gap is the record, not the authority. Found by reading a node's audit log for a refusal whose cause was six hours earlier and unrecorded |
+| **The loopback path removes access without writing an audit entry** | `ops.revoke_member` and `ops.unpin_member` log to the daemon's log and nothing to `audit.db`; the MNP `member_revoke` handler doing the same work audits it (and `member_unpin` did, until it left MNP in 6.0). So a removal made from the node page or the CLI — the two doors an operator sitting at their own machine actually uses — leaves the journal showing an admission and then, whenever that person next connects, an `auth_failed` ("not admitted by the roster") with nothing in between to explain it. §5.4's signed transcript is not what is missing: a loopback caller is authorized by being on localhost with the run token and signs nothing, so the gap is the record, not the authority. Found by reading a node's audit log for a refusal whose cause was six hours earlier and unrecorded |
| **The Create group wizard calls two hooks after an early return** | `CreateGroupWizard` (`create-group-page.js`) returns during node detection, before its `useRef`/`useEffect` for provisioning, so the hook count changes between renders. Preact tolerates a list that grows, and nothing is known to break; `test_hook_ordering.py` checks declaration order, not this. Found while tracing the frozen-fields report, which had another cause (`ask.js`) |
| **A node key is read from the terminal or the desktop client, never a browser** | **Accepted.** `meshbay-node status` on the node's own machine and Node → Overview in the desktop client are the two places the key can be read; the Node page is Electron-only, because `platform.node` resolves to "not available" without the bridge, and no hub route exposes the key. The create-group wizard links it automatically over that same bridge, so the manual paste in **Profile → Link Node** exists for the operator who runs the node from a terminal and the hub from a browser — who has a terminal by definition. Anyone linking a node is already at a shell prompt, so a browser-reachable copy would buy nothing and widen what the hub knows about the node |
| **Listing a group's folders walks every root on the event loop** | `index_sync_message` (`transport/wire.py`) builds its `dirs` field with `list_dirs`, an `rglob("*")` over every root, and nothing sends it off the loop: the WebRTC `index_sync` handler, the daemon's index push and QUIC all call it inline. So each index request from any member is a directory walk of the whole library that every other peer on the node waits behind. `test_disk_io_off_loop.py` never saw it, because it reads the transport's own modules and the walk is one call away in `wire.py`. Found by widening what that test reads, not by a symptom |
| **A loopback eject or plug reaches open pages late** | `ops.eject_root` and `ops.plug_root` flip the live set and tell nobody; over MNP the broadcast `root_eject_ack` / `root_plug_ack` is what moves every open table. So an eject made from the desktop application or the CLI shows on members' pages only with the next index push — for a plug, the end of its rescan; for an eject, whatever changes next. `ops.update_root` had the same silence and now calls `DirectoryIndexer.publish_roots`; the same call belongs in these two. Found while moving the `writable`/`removable` switches to the loopback door (MNP 6.0) |
+| **A group key rotated from the Node page leaves the chat key where it was** | `ops.set_gek(rotate=True)` replaces the group key and opens no chat epoch; the MNP `gek_rotate` handler opened one itself (`_new_chat_epoch`), and it was the only door that did — but no client ever sent it, and it is gone since 6.0. The removals that matter (revoke, unpin, device revoke) open an epoch in `ops` for every door, so what is missing is the follow-through for an operator who rotates by hand: §4.5's "rotate after a removal" means it for chat too. The fix is the `_after_removal` shape — `open_chat_epoch` inside `ops.set_gek` when `rotated`. Found while removing the MNP message |
| **The transcoded-seek test passes without transcoding** | `test_a_transcoded_video_keeps_accurate_seeking` (`test_stream_seek_audio_alignment.py`) forces the re-encode branch by swapping the module's `BROWSER_INCOMPATIBLE_VIDEO_CODECS`, then checks only that the result has no audio gap. The copy path also leaves no gap on that clip, so pointing the swap at a module the streaming code does not read still passes: the test cannot tell that the branch it is named after never ran. It should assert the re-encode happened (the `re-encoding` log line, or the encoder in the ffmpeg argv). Found by breaking the swap on purpose while moving the streaming code |
---