diff options
Diffstat (limited to 'docs/MESHBAY_NODE_PROTOCOL.md')
| -rw-r--r-- | docs/MESHBAY_NODE_PROTOCOL.md | 104 |
1 files changed, 78 insertions, 26 deletions
diff --git a/docs/MESHBAY_NODE_PROTOCOL.md b/docs/MESHBAY_NODE_PROTOCOL.md index ce99a40..9e51dc9 100644 --- a/docs/MESHBAY_NODE_PROTOCOL.md +++ b/docs/MESHBAY_NODE_PROTOCOL.md @@ -413,7 +413,8 @@ fails if a transport skips a step. |-------------------------------------------------------------->| | 8. rebuild binding; | | refuse if empty; | - | compare_digest(proof) | + | compare_digest(proof); | + | roster admits the user | | 9. session authenticated: | | frame limit -> 64 MiB, | | peer registry, audit | @@ -480,6 +481,7 @@ absent. `verify_proof` compares with `hmac.compare_digest`. | not on the denylist for `user_id`, `jti` **or** `group_id` | `Token revoked` | all three targets, and persisted to disk: a revocation that a restart forgets is not one | | `group_id ∈ token.groups` | `Not a member of this group`, code `not_a_member` | the membership check itself — a token is proof of an account, never of a group | | `group_id ∈ node.hosted_groups` | `Group not hosted on this node`, code `not_hosted` | the hub may hand a client several nodes for one group, and only some of them host it | +| *(after the proof)* the roster admits `sub` for `group_id` — an active member row, or the node-wide operator row | `This node has not admitted you to this group`, code `not_authorized_for_group` | the key proves possession and the token the hub's view; the node's own answer is the roster. Someone revoked here but still a hub member, holding the key, is refused a session | `AuthorizedPeer` carries `user_id`, `group_id`, `username`, `jti` — and deliberately **no user public key**. `username` is read from a `username` claim that neither the MNP @@ -489,13 +491,17 @@ would let whoever issues tokens decide it instead. Identity keys are pinned by t node's roster. The hub certifies accounts, not keys. `not_a_member` means the hub did not count this account a member of the group when it -minted the token. The MNP token is minted for each connection, from the membership the -hub holds at that moment, so a stale `groups` claim is no longer the usual cause; the +minted the token. The MNP token is minted for each connection and names **only** that +connection's group (`POST /v1/nodes/mnp-token {node_pk, group_id}`; `groups` is empty +for a non-member), so the operator it is handed to learns nothing of the member's +other groups. A stale `groups` claim is therefore not the usual cause; the client still refreshes its session once and retries on that code before telling someone who was just invited that they are not a member. `not_hosted` is the client's signal to try the **next** node the hub offered for the -group rather than to report a failure. `/v1/groups/{id}/nodes` returns every node +group rather than to report a failure. The hub offers only nodes whose account owns +the group or that the group's owner approved as hosts — a member's node holds the +group key and would pass this handshake like the real host. `/v1/groups/{id}/nodes` returns every node registered for the group, in hub registration order, and that order is not a ranking: a node listed first is not necessarily one that holds the group's files. Refusing without a code made this indistinguishable from a refusal the reader has to act on, @@ -620,7 +626,7 @@ immediately after these three, so the table below is the window, exhaustively. |---|---|---| | `keypair_bundle_fetch` | the client's own identity keys for this node live in an encrypted bundle stored on it | counts against `MAX_PRE_PROOF_FETCHES` = 4; audited | | `gek_bundle_fetch` | the wrapped group key is what the proof is computed with | same counter | -| `join_request` | a first-time member holds no group key at all. Accepted after the proof as well — an operator pairing a browser is already connected — because its authority comes from the pairing code and the signature, never from the session state | 5 attempts per connection, 20 failures per 600 s node-wide | +| `join_request` | a first-time member holds no group key at all. Accepted after the proof as well — an operator pairing a browser is already connected — because its authority comes from the pairing code and the signature, never from the session state | 5 attempts per connection; wrong codes: 5 per account and 20 node-wide per 600 s, consulted only when a code is tried | Device linking (§9) is **not** in this window. `device_add_request` and every message after it are answered only on an authenticated session, and the device budget of 5 @@ -630,10 +636,12 @@ Exceeding the fetch budget is audited as `pre-proof fetch flood` and answered `Too many requests`. Every fetch in this window is written to the audit log with the message type, because this is a disclosure surface a hub that forges a JWT can reach: the hub mints the tokens, so it can present one for any account, and what it can then -ask for is that account's *encrypted* keypair bundle. The bundle is useless without the -account passphrase, which is why the window is bounded and audited rather than closed -— and it closes for good when clients stop storing keypair bundles on other people's -nodes. +ask for is that account's *sealed* keypair bundle. The bundle opens only with the +account passphrase and the account's pepper (`MESHBAY_DESIGN.md` §3.7) — and the hub +holds the pepper, so this is the one place where a hub can still search a passphrase +offline, which is why the window is bounded and audited rather than closed. It is +empty for an account whose browser access is off: the desktop application keeps its +identities and leaves no bundle on any node. ### 7.1 Identity bundles @@ -643,7 +651,7 @@ nodes. |<- keypair_bundle_resp {v, found, | | [bundle_enc], [bundle_enc_recovery]} ----------| | | - | decrypt bundle_enc with the passphrase-derived bundle key, + | open bundle_enc with this node's bundle key (K_node), | or bundle_enc_recovery with the recovery key | | |-- keypair_bundle_store {v, bundle_enc, | after minting or re-wrapping @@ -653,26 +661,52 @@ nodes. |-- keypair_bundle_delete {v} ---------------------->| withdraw the backup ``` -* The bundle is opaque to the node: it is encrypted client-side under a key derived - from the account passphrase (`keyderive.js`), and optionally a second copy under the - account recovery key. The node stores bytes and serves them back to the same - `user_id`. +The node stores both fields as it receives them. What the client puts in them is: + +``` +bundle_enc = base64( "MBK3" ‖ pepper version (1 byte) ‖ nonce (12) ‖ + AES-256-GCM(K_node, JSON {skEd, skX}, aad) ) +aad = "meshbay:bundle:v3|" + account id + "|" + node public key (base64) +K_node = HKDF-SHA256(M, info = "meshbay:bundle:v3|node|" + node public key) +M = HKDF-SHA256(A ‖ pepper, info = "meshbay:bundle-master:v3|" + account id) +``` + +`bundle_enc_recovery` has the same form under the recovery key, with pepper version +0. `skEd` and `skX` are PKCS#8, base64. The node public key is the one the node +proved in its signed challenge (§5), so a bundle is sealed for — and opens only on — +the node that proved it, for the account that stored it. + +* The bundle is opaque to the node: it is sealed client-side as above + (`keyderive.js` in a browser, `keyring.js` in the desktop application's main + process), and optionally a second copy under the account recovery key. The node + stores bytes and serves them back to the same `user_id`. +* **A client reads `MBK3` and nothing else.** A bundle in an earlier format was + sealed under the passphrase alone; the client refuses it by name + (`bundle_format_retired`) and does not mint a replacement identity, which would + leave the node pinning a key nobody holds. `member unpin` drops the bundle, and + the next join is a first contact. * `keypair_bundle_store` is accepted **after** authentication (it is not in the pre-proof list); the fetch is what happens before. * A `store` omitting `bundle_enc_recovery` leaves any existing recovery copy in place. -* Identity keys are **per node**. There is nothing to carry between nodes, and an - operator who cracks the copy on their own disk gets a key that opens nothing +* Identity keys are **per node**, and so are bundle keys. There is nothing to carry + between nodes, and an identity or a `K_node` taken from one node opens nothing anywhere else. -* `keypair_bundle_delete` is **reserved for `device_policy`** (`MESHBAY_DESIGN.md` - §3.7, open item O3): the node honours it, and no interface sends it yet. Withdrawing - the bundle is only safe once the account has chosen not to need it from a browser — - a lone button would strand the next browser that signs in. +* `keypair_bundle_delete` is sent by the desktop application for an account whose + browser access is off (`MESHBAY_DESIGN.md` §3.7): after connecting, it withdraws any + bundle a node still holds for the account, and stores none. With browser access on, + it stores the identity it holds, sealed as above, and seals it again when the + account's `M` has changed since — which is how a passphrase change reaches a node + that was offline when it happened. Browser access is decided in the application, + never on the wire: nodes do not read it. ### 7.1a Per-account blobs (MNP 3.1) The same shape as a keypair bundle with a different payload — playlists today (`docs/playlists.md` §8). The node stores bytes it cannot read for an account it -already holds a bundle for, so this adds **no new trust boundary**. +already admits, so this adds **no new trust boundary**. The client seals them under +`K_pl = HKDF-SHA256(M, info = "meshbay:playlists:v2")` — the one key every node of +the account shares, since a playlist is read from any of them — and a blob that does +not open under it is overwritten with the client's own copy, never treated as newer. ``` C N @@ -822,7 +856,6 @@ Evaluated in order (`_do_join_request`): | Condition | Outcome | |---|---| | `join_attempts >= 5` on this connection | `error: Too many attempts` | -| `>= 20` node-wide failures in 600 s | `error: Pairing temporarily locked`, audited `join_throttled` | | key not 32 raw bytes, or bad base64 | `join_result{ok:false, reason:"invalid_keys"}` | | `\|ts - now\| > 120` | `stale_request` | | `group_id` non-empty and != session group | `group_mismatch` | @@ -831,6 +864,7 @@ Evaluated in order (`_do_join_request`): | account has devices here, this key is not one | `unknown_device` — the way in is a device-add (§9), not a new invite | | device known, no member row, group policy `open` | member row created (`approved_by: "open-join"`) | | device known, a **pending invite** exists for this user | code required even for a known device; `code_required` / `code_invalid` on failure | +| a code is about to be tried and this account has `>= 5`, or the node `>= 20`, wrong codes in 600 s | `error: Pairing temporarily locked`, audited `join_throttled`. Only `code_invalid` counts, and nothing else is gated by it: every member reconnecting gets the key through this message, so a lock applied before recognition would let one member refuse it to everybody | | device known, not an active member of the session group, a code offered | redeemed like any code — an invitation **link** reaches here from someone pinned through another group, or removed and invited back; `code_invalid` on failure | | device known, member row resolved | `join_result{ok, recognised:true, role}` + wrapped GEK | | unknown device, no code, policy `open` | pin TOFU, admit, wrap (`via: "tofu"`, audited) | @@ -967,6 +1001,9 @@ D_req = "meshbay:device_req:v1" || LP(node_pk) || LP(user_id) || LP(pk_ed) || D_add = "meshbay:device_add:v1" || LP(node_pk) || LP(user_id) || LP(pk_ed) || LP(pk_x) || LP(nonce_s) || LP(ts) signed by a PINNED device +D_rev = "meshbay:device_revoke:v1" || LP(node_pk) || LP(user_id) || LP(pk_ed) || + LP(nonce_s) || LP(ts) signed by a PINNED device + code_hash = sha256( code "\x1f" pk_ed25519_b64 "\x1f" pk_x25519_b64 ) ``` @@ -975,6 +1012,11 @@ code_hash = sha256( code "\x1f" pk_ed25519_b64 "\x1f" pk_x25519_b64 ) * `D_add` deliberately **omits the code**: the code is a bearer secret used to find the request, never signed, never echoed. What is signed is the key pair being admitted, so a signature collected for one device cannot admit another. +* `device_add` **must name a pending request** (`code_hash`) filed by the same + `pk_ed25519` and `pk_x25519`; without one, or with one filed by other keys, it is + refused and the request is not spent. A countersignature alone admits nothing. +* `D_rev` has a prefix of its own. Were a retirement signed over `D_add`, a signature + given to retire a key would admit that key wherever it is not pinned yet. * Because both keys go into `code_hash`, a node cannot answer the approver with a substituted key: the approver recomputes the hash from what it typed and what it was given. Nothing here rests on a human comparing digits. @@ -988,7 +1030,7 @@ code_hash = sha256( code "\x1f" pk_ed25519_b64 "\x1f" pk_x25519_b64 ) N -> C device_list_result {pending, devices: [{pk_ed25519, label, pinned_at, pinned_via, added_by_pk, is_this_one}]} - C -> N device_revoke {pk_ed25519, ts, sig over D_add for the victim's keys} + C -> N device_revoke {pk_ed25519, ts, sig over D_rev for the victim's key} N -> C device_add_ack {revoked: pk_ed25519} ``` @@ -2210,6 +2252,11 @@ version can no longer connect" instead. An unreachable hub is deliberately *not* as too old: a captive portal or a closed laptop must not make starting the application impossible. +`client.minimum` also moves for a change the protocol does not see. The keypair bundle +is opaque to the node, so the `MBK3` format (§7.1) moved no protocol version — but a +desktop client older than 0.17.0 still writes the format it replaced, so 0.17.0 is the +minimum. + **The version a peer announces is only as good as the number it ships with.** Every package in the tree carries one version, and a test fails if two disagree — a client announcing a number from a different scheme sorts wherever that scheme puts it, and @@ -2292,8 +2339,9 @@ walks through the gate meant to stop it. filename and no path anywhere in them (§11.1a). * **Rotation is the only thing that removes access.** Revoking a member stops the node serving the next key; the current key and anything already downloaded stay readable. - A chat epoch is opened at the same time, which stops them reading what is said next — - not what was said before, which they could already read. + A chat epoch is opened at the same time, and their connections are closed and refused + from then on (the handshake consults the roster), which stops them reading what is + said next — not what was said before, which they could already read. --- @@ -2321,6 +2369,10 @@ device add "meshbay:device_add:v1" LP(node_pk) LP(user_id) LP(pk_ed25519) LP LP(nonce_s) LP(ts) -> Ed25519 by an ALREADY-PINNED device of the same account +device rev "meshbay:device_revoke:v1" LP(node_pk) LP(user_id) LP(pk_ed25519) + LP(nonce_s) LP(ts) + -> Ed25519 by an ALREADY-PINNED device of the same account + device hello "meshbay:device_hello:v1" LP(node_pk) LP(group_id) LP(user_id) LP(pk_ed25519) LP(nonce_s) LP(ts) -> Ed25519 by the device claiming this connection @@ -2366,7 +2418,7 @@ LP(x) = uint32be(len(x)) || x every field, no exceptions | Code lifetimes (default, settable) | invitation 7 d, operator pairing 24 h, device request 1 h | `roster.py` | | `MAX_PRE_PROOF_FETCHES` | 4 per connection | `webrtc/dispatch.py` | | `MAX_JOIN_ATTEMPTS` | 5 per connection | `webrtc/admission.py` | -| `MAX_JOIN_FAILURES_WINDOW` / `JOIN_FAILURE_WINDOW` | 20 / 600 s, node-wide | ” | +| `MAX_JOIN_FAILURES_PER_ACCOUNT` / `MAX_JOIN_FAILURES_WINDOW` / `JOIN_FAILURE_WINDOW` | 5 per account / 20 node-wide / 600 s, wrong codes only | ” | | Device attempts | 5 per connection | ” | | `MAX_DEVICES_PER_USER` | 5 | `roster.py` | | `MAX_LINK_INVITES_PER_GROUP` | 20 unredeemed invitation links | `roster.py` | |