aboutsummaryrefslogtreecommitdiffstats
path: root/docs
diff options
context:
space:
mode:
Diffstat (limited to 'docs')
-rw-r--r--docs/MESHBAY_DESIGN.md224
-rw-r--r--docs/MESHBAY_NODE_PROTOCOL.md236
-rw-r--r--docs/playlists.md5
3 files changed, 311 insertions, 154 deletions
diff --git a/docs/MESHBAY_DESIGN.md b/docs/MESHBAY_DESIGN.md
index 78376a3..2efedca 100644
--- a/docs/MESHBAY_DESIGN.md
+++ b/docs/MESHBAY_DESIGN.md
@@ -16,7 +16,7 @@
> them — it names the invariant that holds today, not the incident that produced
> it. §13 is the register of those labels.
>
-> Wire versions at the time of writing: **MNP 4.0** (oldest peer accepted 4.0),
+> Wire versions at the time of writing: **MNP 5.0** (oldest peer accepted 4.0),
> **MHP 0.1**, packages **0.16.0**. The normative source for the wire format is
> `MESHBAY_NODE_PROTOCOL.md`; this document states the design the protocol
> serves, not its byte layout.
@@ -99,7 +99,7 @@ opens them.
│ └─────────┘
MHP 0.1 │ signalling (SDP/ICE, <1 KB), presence, revocation push MHP
│
- ┌────┴────┐ MNP 3.0 ┌──────────┐
+ ┌────┴────┐ MNP 5.0 ┌──────────┐
│ node │◄──────── WebRTC DataChannel / QUIC ──────────►│ client │
└─────────┘ index, file chunks, streams, chat, admin └──────────┘
holds the files browser SPA or desktop
@@ -128,8 +128,9 @@ else about content.
This is decision **E9**, and it is the rule any new feature is measured against.
A feature that wants a row on the hub about a group's content is a feature that
has misunderstood the model. It has been re-verified at each content-model change:
-`SwarmSource` carries a content hash, a node id and an endpoint — **no paths, no
-filenames** — and private groups register nothing at all (**H7**).
+**no node registers a content hash with the hub**, for any group (**H7**); the only
+hashes the hub holds are those of public content somebody reported (§7.5) — **no
+paths, no filenames**.
---
@@ -293,7 +294,7 @@ comes from the hash binding:
|---|---|
| Request TTL | 1 h, `[node] device_request_ttl_minutes` |
| Devices per account per node | 5 |
-| Attempts per connection | 5, then a node-wide lockout |
+| Attempts per connection | 5, audited on exhaustion |
| Filing a request | requires the account to have at least one pinned identity already |
`member unpin <user>` removes **every** device of an account, and is the only
@@ -489,9 +490,10 @@ The Members tab can ask the hub to mail the code to the invitee's address on fil
can join in the invitee's place. It is offered because a code that arrives on its
own is worth more to most groups than the property, and it is stated rather than
hidden: the box reads *"Send the invitation by e-mail (may land in spam)"*, is
-checked by default, and is **remembered per account** (the `invite_email`
-preference), so an operator who unticks it once is not asked to again. Unticked,
-the hub is never called and the table above holds exactly. The CLI mails nothing.
+**unticked by default** — giving the hub the code is something the inviter opts into,
+never something done unasked — and is **remembered per account** (the `invite_email`
+preference), so an operator who ticks it once is not asked to again. Unticked, the
+hub is never called and the table above holds exactly. The CLI mails nothing.
The same box sits under the link form, sharing the same preference. Ticked, and
with an address typed, the hub mails the link to that address — and
@@ -519,8 +521,7 @@ group key.** That is a property of open joining, not a defect of this design.
Content in such a group is protected from the network and from non-members, and
from nobody else.
-Note the axis. **`visibility`** (public/private) controls discoverability and swarm
-hash registration (**H7**). **`join_policy`** (open/request/invite) controls
+Note the axis. **`visibility`** (public/private) controls discoverability. **`join_policy`** (open/request/invite) controls
admission. Only the second decides whether a code is required: a public group with
`join_policy = "invite"` keeps the code, because being findable is not being open.
@@ -549,8 +550,10 @@ rendered as a grouped mnemonic. `recovery_key = HKDF-SHA256(R, info =
"meshbay:recovery:v1:" + username)` — HKDF and not Argon2, because `R` has 256 bits
and there is nothing to brute-force. Every time an identity bundle is written to a
node, a **second copy** is written beside it wrapped under `recovery_key`
-(`bundle_enc_recovery`, additive on the wire). `R` is a pass-through: offered in
-the registration email by default, never written to any database, never logged.
+(`bundle_enc_recovery`, additive on the wire). `R` is a pass-through: shown once
+on screen at registration and mailed with the verification code only if the person
+ticks the box for it (unticked by default), never written to any database, never
+logged.
The reset endpoints are built to leak nothing. `POST /v1/users/password/reset-request`
requires **the username and the email on file as a pair**, checked against a blind
@@ -615,7 +618,8 @@ control. It closes for a native device unconditionally, because that device's ke
is in no bundle anywhere. It closes for an *account* only when no browser needs a
bundle on that node — which needs `device_policy {allow_bundle: false}`, **signed
by a pinned key** so the decision is the user's and never the hub's (open item
-O3).
+O3). Withdrawing a bundle already exists on the wire (`keypair_bundle_delete`) and
+is not offered in the interface: it belongs with that decision, not before it.
---
@@ -635,17 +639,17 @@ Every private key lives in an encrypted keystore on the machine that owns it. Th
hub never sees one.
**Domain separation is consistent and mandatory.** Every derivation uses a
-distinct `info` string, and the AES variant adds an `:aes` suffix so two ciphers
-can never derive the same key from one group key. This is a small detail that
-prevents cross-protocol key reuse, and it is checked rather than assumed.
+distinct `info` string, so no two purposes can derive the same key from one group
+key. The chunk and wrap strings end in `:aes`, left from a second cipher that no
+longer exists; it stays because it is part of every key already derived.
### 4.2 Group key wrapping (ECIES)
```
wrap: sk_eph, pk_eph = X25519.generate() # fresh per bundle
shared = X25519(sk_eph, pk_recipient)
- wrap_key = HKDF(shared, salt=pk_eph, info="meshbay:gek_wrap:v1", len=32)
- wrapped = AEAD(wrap_key).encrypt(nonce, gek, aad=pk_recipient)
+ wrap_key = HKDF(shared, salt=pk_eph, info="meshbay:gek_wrap:v1:aes", len=32)
+ wrapped = AES-256-GCM(wrap_key).encrypt(nonce, gek, aad=pk_recipient)
bundle = pk_eph ‖ nonce ‖ wrapped
unwrap: shared = X25519(sk_recipient, pk_eph) # same derivation
@@ -675,20 +679,19 @@ time. This avoids double storage and makes key rotation feasible without
re-encrypting terabytes.
```
-disk (plaintext) → compress → per-chunk AEAD under a group-derived key → transport → client
+disk (plaintext) → per-chunk AES-256-GCM under a group-derived key → transport → client
```
- Chunk size 1 MB: amortises AEAD overhead and enables seeking, because each chunk
is independently decryptable.
-- `chunk_key = HKDF(GEK, salt=None, info="file:" ‖ blake3(file) ‖ ":chunk:" ‖ index)`.
+- `chunk_key = HKDF(GEK, salt=None, info="file:" ‖ blake3(file) ‖ ":chunk:" ‖ index ‖ ":aes")`.
The salt is omitted deliberately: the group key is CSPRNG output and already
uniform, so the file and chunk context belongs in `info`, which is the correct
HKDF usage (**M5**, first review).
- **Chunk authentication is the AEAD tag**, not a per-chunk signature. The tag
authenticates the ciphertext under a key only members hold, which is what the
signature was for.
-- Compression precedes encryption, because compression is ineffective on
- ciphertext.
+- **Chunks are not compressed**: a chunk is encrypted and sent as it was read.
- Upload chunk size is 48 KB, which is what fits the SCTP limit after msgpack
overhead.
@@ -929,7 +932,7 @@ implementations of one security check is **C6** waiting to happen.
```
client → node handshake {token, group_id, nonce_c, v, v_min}
-node authorize_token() JWT · scope · denylist · group_id · membership · hosting
+node authorize_token() JWT · aud · scope · node · denylist · group_id · membership · hosting
node → client handshake_challenge {nonce_s, node_pk, sig} sig: Ed25519 over the challenge (3.4)
── pre-proof window: bundle fetch, join ──
client → node handshake_response {proof}
@@ -956,7 +959,9 @@ buys, per the convention at the top: a client that knows which node it means to
reach can refuse to send a code anywhere else — against a hijacked signaling path
and against a second host of the same group. It proves *a* key, not the *right*
one: it helps only a client that already knows which key to expect. A wrong
-signature is refused; an absent one is an older node, discovered from its answer.
+signature is refused, and so is an absent one: every node the floor admits signs
+whenever it has a channel binding, and one without a binding could not complete the
+proof anyway.
**Channel binding is mandatory and an absent one is refused** — never degraded to
nonce-only, which would silently drop MitM detection:
@@ -1000,7 +1005,9 @@ for `hubFetch` and signaling, and a short-lived **MNP token** (`aud = MNP_AUD`,
from `POST /v1/nodes/mnp-token`) that carries the member's `sub`, `groups` and
`jti` and is the only thing presented in the handshake. The node binds `MNP_AUD`
when it decodes, so a session token is refused here; the hub API binds its own
-audience, so an MNP token captured by an operator is refused there. It is
+audience and requires `exp`, `sub` and a known `scope`, so an MNP token captured
+by an operator is refused there, and so is any other hub-signed token that is not a
+session (a revocation broadcast, an MHP token). It is
checked once, before the proof, so its short life never interrupts a transfer or
a film already playing — a reconnect fetches a fresh one. `meshbay_common/tokens.py`
holds the two audience strings, shared by the hub that issues and the node that
@@ -1031,10 +1038,11 @@ requester and it therefore grants nothing across accounts.
`gek_required: false` bypass (**NS8**).
**Refusals carry a code**, not only a sentence, because a client can act on a code.
-`not_a_member` in particular is usually a token issued before the person was added
-to the group — `groups` is baked in at sign-in and the hub pushes no updates — so
-the client refreshes once and retries rather than telling someone who was invited a
-minute ago that they are not a member.
+`not_a_member` means the hub did not count the account a member when it minted the
+token; since the MNP token is minted per connection from the membership the hub holds
+then, a stale `groups` claim is no longer the usual cause. The client still refreshes
+once and retries on that code before telling someone who was invited a minute ago
+that they are not a member.
### 5.3 Correlation and liveness
@@ -1063,7 +1071,15 @@ TTL 120 s. **The client reconstructs the transcript from announced fields and
refuses to sign if the operation or subject is not what the user asked for**
(**H5**) — a challenge of opaque random bytes signed blind is an unbound signing
oracle. The transcript's subject names the *outcome*, not the operation: what the
-operator is shown before signing has to be what happens.
+operator is shown before signing has to be what happens. **Every value the node acts
+on is in the subject**: the signature covers nothing else of the request, so an
+operation whose effect is several values signs all of them as canonical JSON — a
+root's path *and* whether every member may write there, a group's name *and* the
+directory it exposes — and a secret by its SHA-256, since the subject is audited.
+
+What waits for a signature is bounded: anyone authenticated can ask for a challenge,
+so a connection holds at most eight pending, each at most 64 KiB, expired ones
+dropped (**AV31**).
Verification is against `roster.operator_pks()`, rebuilt from node state, **never**
from anything in the response.
@@ -1271,19 +1287,21 @@ new client, and the hub, node and SPA deploy together.
with each MAJOR, so `check_version` refuses at the handshake any peer that cannot
meet one: an upload is sealed or it is not sent; a transfer has a real lease or it
does not run; there is one app-directories op and no wrappers behind it. The client
-records the version its peer declared, for diagnostics, and **branches on none of
-it**.
+checks the version its peer declared and **branches on none of it**.
> A capability flag on a peer whose floor already guarantees the capability is a
> branch that can only ever take one path — until somebody lowers the floor, at
> which point it silently takes the other. **A field kept "just in case" is how
> the branches come back.**
-**The floor is not the current version, and MINOR additions are why.** It is
-`MNP_MIN_SUPPORTED` in `handshake.py`, it equals the last MAJOR, and 3.1, 3.2,
-3.3 and 3.4 have all been added above it without moving it. So a peer can be reachable and
-still not do something the current version can, and the client has to cope with
-that — **by reading the peer's own answer, never by comparing version numbers**.
+**The floor is not necessarily the current version, and what is added above it is why.**
+It is `MNP_MIN_SUPPORTED` in `handshake.py`, and it is the last MAJOR that had to
+refuse at the handshake: 3.1–3.4 were added above the 3.0 floor without moving it,
+and 5.0, a MAJOR confined to four signed operations that a peer across the break
+refuses to sign, sits above the 4.0 floor. So a
+peer can be reachable and still not do something the current version can, and the
+client has to cope with that — **by reading the peer's own answer, never by comparing
+version numbers**.
3.2's audio tracks are the worked example: the node lists them in `stream_init`,
the client draws its selector from that list, and a node that sends no list gets no
selector. 3.3's subtitles repeat it exactly, and add the case where the list is
@@ -1813,8 +1831,10 @@ the next node on a `not_hosted` refusal (`MESHBAY_NODE_PROTOCOL.md` §6.3).
The node authenticates to the hub with an Ed25519 signature over a domain-separated
timestamped message — **no password and no auth key on a node** — and receives a
-`scope: "node"` token that is refused for group management. The operator manages
-groups from a client (**NS7**).
+`scope: "node"` token that is refused for group management and on every admin and
+moderator route, even when the account behind it holds a hub role: what a node may do
+is its operator's roster pin, and a hub role is a person's, exercised from a client.
+The operator manages groups from a client (**NS7**).
Signaling is rate-limited, SDP-size bounded, capped per user, and **the caller must
share an active group with the target node**. Otherwise any authenticated user
@@ -1930,8 +1950,9 @@ listed with what they check. The hub publishes no API description — no `/docs`
`/redoc` or `/openapi.json` — in the code, not in a proxy rule, so a packaged
install behind any proxy publishes none either.
-The mail bounds (`mail.*`) and the sign-in lockout (`login.max_failures`,
-default 4, and `login.lockout_minutes`, default 60 — §7.7) live in the same table
+The mail bounds (`mail.*`), the report policy (`reports.*`, §7.5) and the sign-in
+lockout (`login.max_failures`, default 4, and `login.lockout_minutes`, default 60 —
+§7.7) live in the same table
for the same reason: they are what an operator changes while the hub is serving,
from the panel, without a restart. Each value is clamped to published bounds, and
`max_failures = 0` turns the lockout off.
@@ -1950,17 +1971,71 @@ The client shows the real state, not a blanket one. **Revocation is honoured by
nodes** and the denylist survives a restart (**H4**); signaling refuses a group that
is not active.
+**Revocation has one door, and it broadcasts.** Only an administrator revokes, and
+only through `POST /v1/admin/revoke`, which signs the revocation and pushes it to every
+connected node — the same signed broadcast an administrator's account deletion sends
+(§7.7). The user and group PATCH handlers refuse `revoked` outright, because a
+status written there reached no node and behaved as a suspension while claiming to be
+a revocation. Moving a group or an account *out* of `revoked` is an administrator's
+call too, and it changes the hub row only: the nodes keep enforcing the revocation
+they received, so the hub and the nodes then disagree until each operator clears it
+(`denylist clear`). That is why the table says "no".
+
**Moderator is not administrator.** The user-patch handler is split by field: a
moderator may act on the fields moderation needs and may not write `role`.
-**Content reporting requires authentication, distinct reporters and a rate limit,
-and is refused when public groups are off.** An unauthenticated endpoint that
-blocklists a content hash after two reports is a network-wide censorship and DoS
-primitive for anyone who learns a public file's id.
+**A file is reported by a member of the public group it was seen in, and an
+administrator decides.** An unauthenticated endpoint that blocklists a content hash
+after two reports is a network-wide censorship and DoS primitive for anyone who
+learns a public file's id, so every bound here answers what a report costs somebody
+else — a file taken out of a group everyone else uses, and an administrator's time:
+
+- a **person's** account — a node's token is refused — that has existed for
+ `reports.min_account_age_hours` (24 by default);
+- **active membership of the public group** named in the report, answered with one
+ refusal whatever the reason, so the endpoint says nothing about which groups exist
+ or who is in them. The hub cannot check that the file is in that group — it holds
+ no index, by design — only that the reporter could have seen it there;
+- a **daily allowance per account** (`reports.daily_per_account`, 20) besides the
+ per-address rate limit, since an address is one of thousands a subscriber holds;
+ one report per account per hash; a closed list of reasons and a bounded detail;
+- once `reports.review_threshold` distinct accounts (3) have reported a hash, it is
+ **queued for review** and the administrators are notified. An administrator blocks
+ it — which pushes it to the nodes — or dismisses it, and a dismissed hash is not
+ reopened by more reports. The reporter is told the report was recorded and never
+ how close the file is to review.
+
+`reports.auto_block` blocks at the threshold without a review. It is off by
+default and stays an instance's explicit choice, because it makes a handful of
+accounts made for the purpose enough to take a file down. The whole flow is refused
+while public groups are switched off. The interface offers **Report** on a file's
+menu in a public group only.
-The exact-hash CSAM check is **structural, not yet functional** — production
-databases are perceptual — and is stated as such so it is not relied on
-operationally.
+**The content blocklist is applied by the nodes that host a public group, in their
+public groups only.** A node holds the list (`blocklist.py`, persisted beside the
+denylist so a restart while the hub is unreachable does not serve again what had
+stopped being served), fetches the whole of it on every connection to the hub —
+`GET /v1/blocklist`, paged by hash, answered to a node's own token only — and
+receives each addition and removal pushed on its hub socket (`blocklist_update`),
+sent only to nodes registered for a public group. In a public group a blocked file
+leaves the index members are sent, and a request for it, its thumbnail, a stream of
+it, a subtitle track or an audio conversion of it is refused (`content_blocked`);
+the members connected when the list changes are resent the index. Nothing is
+deleted: the file is on the operator's disk, and what they keep is theirs.
+
+Stated per the convention at the top:
+
+- **Private groups are untouched**, by construction: no node sends the hub a
+ content hash (**H7**), so nothing in a private group can be on the list.
+- **Which groups are public is the node's own configuration.** The hub and the
+ node are given the same value when a group is created and the hub never changes
+ it; an operator who edits `node.toml` to call a hub-listed group private takes
+ it out of the list's reach.
+- **It is an exact match on the content id.** A file changed by one byte is
+ another id, and a file past the partial-hash threshold is identified by a sample
+ of its bytes (§6.3). It is a moderation tool, not a guarantee.
+- **The QUIC transport does not apply it** — it is in development and serves no
+ client (§5.1, §15.3).
### 7.6 Federation (MHP)
@@ -2002,7 +2077,7 @@ radius is proof of the passphrase. An admin can delete one too.
The row is **tombstoned rather than dropped**: username released, email and password
hash cleared, node linking key dropped, memberships, notifications, refresh tokens,
-node registrations, device keys and public-swarm sources removed, active tokens
+node registrations and device keys removed, active tokens
refused at once by a status check rather than left to expire. Device keys go
because the desktop client keeps its half: left on the tombstone, the key would
refuse that installation to the next account created from it.
@@ -2067,8 +2142,8 @@ The rules that make this safe:
attempts than the limit. A request that checked no passphrase gives its attempt
back.
- **Every path that checks the passphrase counts on the same row**: sign-in,
- passphrase change and account deletion. A right passphrase clears it; failures
- older than the window age out.
+ passphrase change, changing the e-mail address on file and account deletion. A
+ right passphrase clears it; failures older than the window age out.
- **A lockout refuses passphrase sign-in and nothing else.** Open sessions, token
renewal and device sign-in continue, and a reset code sent to the address on
file clears it — so a stranger who locks a public username costs its owner at
@@ -2462,7 +2537,7 @@ everything registered — a node that predates an application hides nothing.
**An application's directories are the same shape one level down**: one generic
signed op (`app_directories`) keyed by the application's own registry name, stored
-under `<key>_directories`, one MNP message, one loopback route. Adding an
+under `<key>_directories`, one MNP message. Adding an
application adds **no function, no message type and no route** — which is what
"plug-in architecture" has to mean to be worth the phrase.
@@ -2908,6 +2983,15 @@ requirements, not compatibility notes.
| **Watcher reliability** | Change notification drops events under load on Windows, and inotify is unreliable on a FUSE mount. **Periodic reconciliation is mandatory on both platforms** |
| **No symlinks, no POSIX permissions** | Simplifications: nothing to defend against, and the node runs as the user anyway |
+**The client makes a name writable when it saves, and says so.** A node serves
+the name its disk gave the file and never rewrites it — that is the string that
+opens it — so a single file, a zip's entries and the zip's own name are passed
+through `portable-name.js` at the moment of saving (reserved characters become
+`_`, a trailing dot or space goes, a reserved stem gains `_`), and the transfer's
+row names the original when it changed. The rule is `paths.sanitize_for_download`,
+and the two are held byte-identical by a parity test. Two different names can
+still become one — a zip keeps both entries under it.
+
**Case folding is for comparisons the code makes itself** — index identity,
collision reporting, root names, nesting checks. It is *not* needed for the
no-overwrite rule, where the filesystem's own case-insensitive `stat()` already
@@ -2952,7 +3036,7 @@ The bulk of the codebase is portable because the portability rules in §10 were
treated as correctness from the start. What the port needed is registered as
**W1–W9** (§13.7) and is done; packaging is built and awaits a clean-machine run.
-Two Windows-specific design points worth stating here:
+Three Windows-specific design points worth stating here:
- **The node runs in one of three modes, chosen at install and switchable
afterwards** from the Node page: only while the application is open (it starts
@@ -3090,7 +3174,7 @@ be understood, not so the incident can be retold.
| **NS4** | **Operator authority comes from the node's roster and from nowhere else.** No auto-pin from the keystore, no resolution through the hub, no config key — a config naming one is warned about and never obeyed (§3.4, §6.1) |
| **NS5** | The proof is **bound to the transport channel** (DTLS fingerprints / certificate hash), so a signaling relay that substitutes its own cannot produce it (§5.2) |
| **NS6** | **`sender_id` is enforced from the authenticated session, never the wire.** It is what the store keys on; it is not what authenticates a message — the device signature is (§4.5) |
-| **NS7** | The node authenticates to the hub with **Ed25519 and no password**, and its token's scope is refused for group management (§7.2) |
+| **NS7** | The node authenticates to the hub with **Ed25519 and no password**, and its token's scope is refused for group management and on the admin and moderator API (§7.2) |
| **NS8** | **The node refuses connections when it holds no group key.** There is no bypass switch |
### 13.3 Second review (code review) — the default numbering
@@ -3118,7 +3202,7 @@ be understood, not so the incident can be retold.
| **H4** | Revocation reaches nodes, drops live sessions, and **persists across a restart** (§7.5) |
| **H5** | An admin challenge is a **structured, domain-separated transcript naming the operation and subject**, and the client refuses to sign anything that is not what the user asked for (§5.4) |
| **H6** | Unauthenticated work a node will do is bounded: a small pre-handshake buffer, a transcode semaphore, per-user pending-offer caps, and a membership check on signaling (§7.2) |
-| **H7** | **Only public groups register content hashes with the hub.** Private groups register nothing, and the swarm route requires authentication |
+| **H7** | **No node registers a content hash with the hub**, for any group. The only content hashes it holds are of public content somebody reported (§7.5) |
**Medium**
@@ -3161,7 +3245,7 @@ be understood, not so the incident can be retold.
| **M4** *(third review)* | Federation binds a pushed row's source to the signer, checks the token audience, caps the push, rejects replays, and scopes revocation to the peer's own entries (§7.6) |
| **M5** *(third review)* | A CSP and security headers apply to the hub-served application, verified against the running app — a mis-tuned CSP shows as a blank page |
| **M6** *(third review)* | **Withdrawn.** It misread the node registering a hub membership during the CLI invite flow — which is deliberate — as authorization drift |
-| **L1–L11** *(third review)* | Opportunistic hardening: relay-registry proof of possession; delete orphaned modules rather than leaving them to be rewired; decide and document account enumeration; an aggregate upload quota; header-only control-API tokens; a freshness bound on revocation replay; state that the exact-hash content check is structural; validate group-name length and charset; require `exp` and bind an audience on token decode; key the rate limiter through the same client-address helper as everything else; keep diagnostic logging truncated |
+| **L1–L11** *(third review)* | Opportunistic hardening: relay-registry proof of possession; delete orphaned modules rather than leaving them to be rewired; decide and document account enumeration; an aggregate upload quota; header-only control-API tokens; a freshness bound on revocation replay; validate group-name length and charset; require `exp` and bind an audience on token decode; key the rate limiter through the same client-address helper as everything else; keep diagnostic logging truncated |
Two structural recommendations from that review stand as rules:
@@ -3206,15 +3290,12 @@ the wrong thing** — a per-IP rate limit bounds a caller, never the mailbox tha
receives what they cause, which is why `AV10` is a cooldown per *account* under
a limit per IP rather than a tighter limit.
-**Open for decision, not a defect:** `moderation.AUTO_BLOCK_THRESHOLD` is 3.
-Three distinct accounts blocking a hash adds it to the list every node
-enforces, network-wide, automatically, with manual admin removal the only
-undo. That is already far better than the anonymous version it replaced, and
-it is still a censorship primitive an attacker buys for the price of three
-email addresses. Raising it buys little; requiring the reporting accounts to
-be more than a day old would cost a patient attacker a day and cost an honest
-reporter nothing after their first. Left as it is because it is a moderation
-policy rather than a bug, and the person who sets that policy is the operator.
+**Reports lead to a review, not a block** (§7.5). Distinct accounts reaching
+the threshold used to add a hash to the list every node enforces, automatically,
+with manual removal the only undo — a censorship primitive bought for the price of
+three email addresses. Now the reporter must be a day-old account and a member of
+the public group, and the threshold queues the hash for an administrator; blocking
+without review is an instance setting, off by default.
`admin.py` was read under this lens and needed nothing. Its moderator/admin
line is drawn explicitly — a moderator may not change a role, may not revoke,
@@ -3227,9 +3308,9 @@ had already been asked.
| **AV1** | **An empty claim is a claim on nothing.** A node's group set is `authorized ∩ claimed`, and an absent or empty `group_ids` registers it for no group rather than all of its owner's — on registration and on `update_groups` alike (§7.2) |
| **AV2** | **A client treats the hub's node list as candidates, not a ranking**, and tries the next one on a `not_hosted` refusal (`MESHBAY_NODE_PROTOCOL.md` §6.3) |
| **AV3** | **A node speaks only for the groups it is registered for.** `chat_notify` names a group and is checked against that node's set before a notification is written for anyone, and it is rate-limited per node — the fan-out is one write per member |
-| **AV4** | **Nobody names a third party's address.** A swarm source publishes a transport and a port, never a host; where a peer is comes from its node record, stamped with the address its announce arrived from. The number of hashes one account may claim is bounded |
+| **AV4** | **Nobody names a third party's address.** Where a peer is comes from its node record, stamped with the address its announce arrived from |
| **AV5** | **An answer is accepted only from the node the offer was sent to.** A `peer_id` is bound to its node, so no connected node can resolve another's pending offer |
-| **AV6** | **A relay proves possession of its approved key.** A public key is not a password, and the register call is unauthenticated by design — it is not a user — so the proof is the only thing standing between a stranger and where nodes send relayed traffic |
+| **AV6** | **A write that decides where other people's traffic goes proves possession of a key, never presents one.** A public key is not a password. The relay registry this was written for has been removed — nothing called it, and no TURN relay is needed (§11.1) — and the rule stands for whatever replaces it |
| **AV7** | **A node bounds how many peers it holds and how long an unproven one lasts.** The hub's cap is per calling account, which is a limit on each member and not on the machine, so without this an operator's exposure grew with the size of their groups |
| **AV8** | **One account cannot make the hub mail another at will.** The invitation email's subject comes from the group row, never from the request, and the endpoint is metered |
| **AV9** | **No mail is sent from the event loop.** `smtplib` is synchronous and waits up to ten seconds; called from an async handler that wait is the whole instance's, not one request's. Every send goes through `mail.send_off_loop`. **Argon2 is held to the same rule**: every derivation runs on one dedicated worker thread (`auth.*_off_loop`), never on the loop and never two at a time, because two concurrent `lanes=4` derivations deadlock in OpenSSL. **So is the node's disk**: every filesystem call on a group's content — the stat as much as the read, since a stat is what wakes a sleeping disk — goes through `roots.off_disk`, onto one worker thread per root set. A spun-down or network-mounted root answers its first syscall in seconds, and on the loop that is every group, every stream and the hub socket waiting for a platter. **ffmpeg's own output too**, through `asyncio.to_thread` rather than that per-root thread: a temp file is not a group root and has no platter to serialise against, but a whole transcode read inline is still tens of megabytes of blocking read |
@@ -3242,7 +3323,7 @@ had already been asked.
| **AV19** | **Nothing carries the path to the migrations.** `meshbay-hub migrate` derives it from the installed package, so the RPM, the DEB, a venv and a checkout all agree. A unit naming `alembic.ini` names a file whose `%(here)s` stops being true the moment packaging moves it |
| **AV18** | **The hub runs on exactly one worker, and says so at startup.** `_connected_nodes`, `_node_groups`, `_webrtc_answers` and the relay registry are per-process: a second worker makes a node intermittently unreachable for half its members, which is a symptom that describes something else entirely |
| **AV14** | **MHP binds its audience, and the hub reads its own identity at call time.** A token is minted for one peer and accepted by that peer only. `federation.py` bound `_hub_id` and `_hub_sk_pem` at import, which is before `load_hub_keypair` runs, so it signed with `None` and called itself `meshbay.org` whatever the instance was named — and the verifier named no audience for the `aud` the issuer sets, which PyJWT refuses outright. MHP could not complete one authenticated request between two hubs |
-| **AV15** | **A hash is checked for shape before it is a key lookup**, on the unauthenticated blocklist endpoints a node consults |
+| **AV15** | **A hash is checked for shape before it is a key lookup**, on every blocklist endpoint, the administrator's included |
| **AV20** | **Chat is bounded in size and in rate, like every other member-supplied write** (§6.6). A message is a row on the operator's disk that nothing expires, a relayed copy for every connected member and a notification for every member of the group; the only ceiling was the frame size. Uploads had carried four protections and a cap since C5a because somebody asked what one member costs the others on that path, and nobody had asked it on this one |
| **AV21** | **A lease is what the node granted, not what the client called it** (§5.5). `tr` was read as a boolean, so any non-empty string skipped the leaseless ceiling and every cap behind it, and a queued transfer was held back only by the honesty of the client waiting in the queue |
| **AV22** | **The node's own controls take no authority from a hub token** (§6.7). `node_status`, `node_settings_set`, `roster_read`, `denylist_read`, `denylist_clear` and `node_reload` were gated on the account id in the JWT, which is the hub's to choose — NS4 and M3 with the check written the other way round. The gate is a proved operator device, which a hub holding no user keys cannot produce |
@@ -3254,6 +3335,7 @@ had already been asked.
| **AV29** | **An invitation link is bounded on both halves and its mail on the sender** (§3.4, §7.3). Twenty outstanding per group on the node (bearer codes) and on the hub (tickets); and because a link mail reaches an address the hub has no relationship with, at the request of anyone who owns a group, it is counted **per sending account per day** (`mail.invite_link_daily_cap`, 10), under the recipient and instance bounds and outside the recovery reserve (`invite_link` is not a recovery purpose) |
| **AV28** | **How many node keys one account may announce is bounded** (§7.2). Each is a row plus an IP-log row under a one-year retention, so an account in a loop writes a year of storage on the operator's disk having paid only for signatures. Proof of possession (**M8**) settles whose key it is and not how many. Counted only where a row is added: re-announcing a key already held keeps working at the ceiling, or a node that reached it could never refresh its address again |
| **AV30** | **What one member's offers cost a node is bounded per account and per node, and the bound admits the heaviest ordinary account** (§7.2). Each offer makes the node allocate a peer connection. A budget of 120 per node refilled at two a second bounds a member there without touching their other nodes, and it is counted by account because a mobile carrier shares one IPv4 address among many subscribers. Pending offers are capped at 32 per account. Both refusals carry `Retry-After` and the client retries them, because a refused offer otherwise reads as a node that is down |
+| **AV31** | **What waits for a signature is bounded** (§5.4). Any authenticated member can ask for an admin challenge, since the signature is checked afterwards, and a pending challenge kept its whole request until answered — measured, 200 requests of 1 MiB held 400 MiB for the life of one connection. At most eight pending per connection, 64 KiB each, expired ones dropped |
### 13.6 Chat design findings
@@ -3274,7 +3356,7 @@ had already been asked.
|---|---|
| **W1** | Platform directories: no hardcoded XDG paths |
| **W2** | Signal handling is platform-guarded |
-| **W3** | Daemon lifecycle: a per-user startup launcher by default, a scheduled-task service mode offered, switchable after install (§11.2) |
+| **W3** | Daemon lifecycle: three modes — only while the application is open, at sign-in, or as a boot-time scheduled task — chosen at install and switchable after; one start/stop implementation, the CLI's (§11.2) |
| **W4** | Packaging: one per-user installer carrying client and node, with media tools bundled |
| **W5** | File permission calls are skipped where they have no meaning |
| **W6** | Media-tool discovery fails at startup with a stated reason rather than at first use |
@@ -3422,9 +3504,7 @@ process runs it — `systemctl --user` on Linux, Task Scheduler on Windows.
| **A signed upload transcript** | Ownership is recorded by the node and verifiable by nobody else (§5.4). Making it provable is a transcript the uploader signs, stored with the entry — designed in outline, not built |
| Forward secrecy in group chat | **Given up deliberately and on the record** (§4.5). If it becomes a requirement it belongs in 1:1 DM |
| Metadata at the hub | Membership, and who posted in which group and when. A known leak, not a solved problem (§7.1) |
-| The exact-hash content check | Structural, not functional (§7.5) |
-| **QUIC** | Off by default, and **not at parity**: it serves the index and file chunks with no transfer lease, no leaseless ceiling and no root-availability check, does its file I/O on the event loop, and returns exception text to the peer (**L3**). No client speaks it. Either it comes to parity or it goes; until then §5.1's "chat is the only gap" is the one sentence here that overstates the code |
-| **The relay registry** | **Closed in the code**: `relay.RELAYS_ENABLED` is False and every `/v1/relays` route answers 503, as federation does. Nothing in the tree calls them, node or client, and §11.1 measured two ISPs with no TURN relay needed. Kept code that nothing calls is what **L7** says not to keep; it stays only as the proof-of-possession design (**AV6**) until a node needs a relay or it is deleted |
+| **QUIC** | Off by default, and **not at parity**: it serves the index and file chunks with no transfer lease, no leaseless ceiling, no root-availability check and no content blocklist, does its file I/O on the event loop, and returns exception text to the peer (**L3**). No client speaks it. Either it comes to parity or it goes; until then §5.1's "chat is the only gap" is the one sentence here that overstates the code |
| **A very high bitrate wedges the player against a small buffer ceiling** | Where even the *floor* read-ahead does not fit — ninety seconds plus the minute kept behind, at the file's bitrate, above what the engine will hold — the film stalls: measured on the harness at 9.3 Mbit/s against a 100 MB ceiling, 100.8 s of film played in 900 s of wall clock. **Predates the byte budget and is unchanged by it**, to the tenth of a second; what the budget did change there is the refusal count, 1560 → 2. The fix is not a bound at all, it is a second stage of buffer outside the SourceBuffer, which means gating the append path — the riskiest change in this area and not one to make alongside another |
| **The reconnect backoff only wakes on `visibilitychange`** | So a tab that stays visible through an outage — which is what a screen wake lock guarantees while a film is playing — waits out the full backoff, up to 30 s, after the network is already back. Nothing listens for `online` |
| **Per-device revocation has no CLI** | A device is revoked over MNP (`roster.revoke_device`), from a device the node has already pinned. On a headless node the operator's only lever is `member unpin`, which removes **every** device of that account — so the per-device control the roster is built around is reachable from an interface and from nowhere else. §6.7 listed a `meshbay-node member device list\|revoke` verb that was never written, and that listing is how this was found: `USERGUIDE.md` was the first document written by reading the CLI rather than this specification, and the verb it copied out did not run |
diff --git a/docs/MESHBAY_NODE_PROTOCOL.md b/docs/MESHBAY_NODE_PROTOCOL.md
index 8de731e..ce99a40 100644
--- a/docs/MESHBAY_NODE_PROTOCOL.md
+++ b/docs/MESHBAY_NODE_PROTOCOL.md
@@ -1,14 +1,14 @@
# MeshBay Node Protocol (MNP)
-**Wire version:** `4.0` — `meshbay_common/__init__.py` (`MNP_VERSION`)
-**Oldest peer accepted:** `3.0` — `handshake.py` (`MNP_MIN_SUPPORTED`)
+**Wire version:** `5.0` — `meshbay_common/__init__.py` (`MNP_VERSION`)
+**Oldest peer accepted:** `4.0` — `handshake.py` (`MNP_MIN_SUPPORTED`)
**Normative implementation:** `meshbay-common` (`protocol.py`, `handshake.py`,
`groupbox.py`, `chatbox.py`, `adminop.py`, `join.py`, `device.py`, `crypto.py`,
-`webcrypto.py`), `meshbay-node` (`transport/wire.py`, `transport/webrtc_server.py`
+`webcrypto.py`, `tokens.py`), `meshbay-node` (`transport/wire.py`, `transport/webrtc_server.py`
and `transport/webrtc/`, `transport/quic_server.py`, `transfers.py`, `uploads.py`), browser client
(`meshbay-hub/static/transport.js`, `static/crypto.js`).
**Document status:** descriptive specification of the protocol as implemented on
-2026-09-10. It describes the protocol as it stands. Where the code and this document
+2026-09-28. It describes the protocol as it stands. Where the code and this document
disagree, the code is authoritative and this document is the thing to fix.
**It is self-contained.** Every rule below is given with the reason it exists, in
@@ -152,10 +152,14 @@ Codes in use:
| Family | Codes |
|---|---|
-| Handshake | `not_a_member` (§6.3); `version_too_old`, `version_too_new`, `version_unreadable` (§13.1) |
-| Transfers | `transfer_required`, `bad_transfer_id`, `bad_transfer_size`, `bad_kind`, `not_your_transfer`, `too_many_queued` (§11.2) |
-| Upload | `upload_not_sealed`, `no_group_key`, `upload_incomplete`, `bad_chunk_encoding`, `bad_chunk_index`, `invalid_filename`, `no_roots`, `no_such_root`, `no_writable_root`, `root_read_only`, `root_unavailable`, `no_such_directory`, `already_exists`, `not_started`, `too_large` (§11.4) |
+| Handshake | `not_a_member`, `not_hosted`, `wrong_node` (§6.3); `version_too_old`, `version_too_new`, `version_unreadable` (§13.1) |
+| Transfers | `transfer_required`, `lease_not_granted`, `bad_transfer_id`, `bad_transfer_size`, `bad_kind`, `not_your_transfer`, `too_many_queued` (§11.2) |
+| Upload | `upload_not_sealed`, `no_group_key`, `lease_not_granted`, `upload_incomplete`, `bad_chunk_encoding`, `bad_chunk_index`, `invalid_filename`, `no_roots`, `no_such_root`, `no_writable_root`, `root_read_only`, `root_unavailable`, `no_such_directory`, `already_exists`, `not_started`, `too_large` (§11.4) |
| Directories | `root_read_only`, `root_unavailable` (§11.5) |
+| Chat | `chat_too_large`, `chat_rate_limited` (§11.7) |
+| Operator controls | `not_operator`, `too_many_pending`, `too_large` (§10.4) |
+| Metadata | `transcode_not_applicable`, `tmdb_search_rate_limited` (§11.9) |
+| Moderation | `content_blocked` — a file the hub's content blocklist names, in a public group (§11.3) |
Everything else refuses with `detail` alone. A code is added when a client has a
different thing to *do* about the refusal — retry, re-authenticate, offer an update —
@@ -305,7 +309,9 @@ whatever it has (host candidates are enough on a LAN).
| | authorize: shared active group,
| | or an open-join group when public
| | groups are enabled; <=16 KiB SDP;
- | | <=3 pending per user; 30/min
+ | | <=32 pending per account; per
+ | | account and node: burst 120,
+ | | refill 2/s; 600/min per address
| | |
| |--- ws {webrtc_offer, |
| | peer_id, user_id, |
@@ -375,7 +381,11 @@ fails if a transport skips a step.
| else error{code} |
| authorize_token() |
| - EdDSA verify vs hub |
+ | - aud == MNP_AUD, |
+ | exp/sub/scope present|
| - scope == "user" |
+ | - node claim, if set, |
+ | == this node's key |
| - sub non-empty |
| - group_id non-empty |
| - denylist(user,jti,gp)|
@@ -463,8 +473,8 @@ absent. `verify_proof` compares with `hmac.compare_digest`.
|---|---|---|
| JWT verifies under the hub's Ed25519 public key (`EdDSA`) | `Invalid JWT: ...` | |
| `aud == MNP_AUD`, and `exp`/`sub`/`scope` present (MNP 4.0) | `Invalid JWT: ...` | the member presents a short-lived **node-audience** token (`POST /v1/nodes/mnp-token`), not its hub session token — the operator holds whatever is presented, and the session token opens the hub API. The two audience strings are in `meshbay_common/tokens.py` |
-| `node` claim, when set, equals this node's key (MNP 4.0) | `Token is not for this node`, code `wrong_node` | the token names the node it was minted for, so one captured by node A's operator cannot be replayed to node B (E10). A token naming no node is accepted — the hub mints an unbound one only for the requester |
| `scope == "user"` | `Wrong token scope` | a node-scoped daemon token must not be usable as a client token |
+| `node` claim, when set, equals this node's key (MNP 4.0) | `Token is not for this node`, code `wrong_node` | the token names the node it was minted for, so one captured by node A's operator cannot be replayed to node B (E10). A token naming no node is accepted — the hub mints an unbound one only for the requester |
| `sub` non-empty | `Token has no subject` | |
| `group_id` non-empty | `group_id is required` | an absent group means no membership check to make; there is no default group, and a node's first group is not one |
| not on the denylist for `user_id`, `jti` **or** `group_id` | `Token revoked` | all three targets, and persisted to disk: a revocation that a restart forgets is not one |
@@ -472,14 +482,17 @@ absent. `verify_proof` compares with `hmac.compare_digest`.
| `group_id ∈ node.hosted_groups` | `Group not hosted on this node`, code `not_hosted` | the hub may hand a client several nodes for one group, and only some of them host it |
`AuthorizedPeer` carries `user_id`, `group_id`, `username`, `jti` — and deliberately
-**no user public key**. A key arriving in a token would be a key the hub chose, and the
+**no user public key**. `username` is read from a `username` claim that neither the MNP
+token nor the hub session token carries, so it is empty in practice. A key arriving in a token would be a key the hub chose, and the
node records the uploader's key in order to decide who may later delete a file: that
would let whoever issues tokens decide it instead. Identity keys are pinned by the
node's roster. The hub certifies accounts, not keys.
-`not_a_member` is almost always a token minted before the person was added to the
-group (`groups` is baked in at login and the hub pushes no updates), so the client
-refreshes once and retries on that code rather than telling a member they are not one.
+`not_a_member` means the hub did not count this account a member of the group when it
+minted the token. The MNP token is minted for each connection, from the membership the
+hub holds at that moment, so a stale `groups` claim is no longer the usual cause; the
+client still refreshes its session once and retries on that code before telling someone
+who was just invited that they are not a member.
`not_hosted` is the client's signal to try the **next** node the hub offered for the
group rather than to report a failure. `/v1/groups/{id}/nodes` returns every node
@@ -537,10 +550,12 @@ this is the key the client wanted. A client with no expectation learns nothing m
than before, and the ack remains the proof of GEK possession.
Client rule, from the node's answer and never from its version (§13): a `sig` that
-does not verify is a refusal (`Node challenge signature invalid`); no `sig` is an
-older node, whose key is proved only at the ack. A node signs whenever it has a
-binding, and a node with none sends no signature rather than an unbound one — the
-proof would be refused on that connection anyway.
+does not verify is a refusal (`Node challenge signature invalid`), and so is a
+challenge with no `sig` at all. A node signs whenever it has a binding, and a node
+with none sends no signature rather than an unbound one — its proof would be refused
+on that connection anyway, so refusing the unsigned challenge only says so earlier.
+With the floor at 4.0 every peer a client can reach signs, and there is no "older
+node" case to tolerate.
### 6.6 `handshake_ack` fields
@@ -564,10 +579,9 @@ at all — including a decryption.
| `chat_link_preview` | whether the node unfurls links posted here. Absent means on |
| `search_listed` | whether the reader's cross-group Search lists this group. Presentation only — the index is served identically either way. Absent means listed |
| `chat_epoch` | the chat epoch a client must seal under right now (§11.7). There is no `chat_encrypted` beside it, because there is no switch |
-| `transfer_limits` | `{download, upload}` — this member's own caps in this group, so the interface can say "2 of your 2 slots are busy" instead of drawing a bare spinner. Absent reads as "no limit known" and the hint is not drawn; never as "unlimited", which would have the interface contradicting the node (§11.2) |
| `tmdb_enabled`, `musicbrainz_enabled` | per-group metadata lookups |
| `tmdb_token_customized`, `tmdb_language` | node-wide TMDB config; the token itself is never sent |
-| `indexing` | `{scanning, scanned_bytes, total_bytes}` so a client connecting mid-scan shows progress immediately. Never a path or filename |
+| `indexing` | the `index_progress` counters (§11.1) so a client connecting mid-scan shows progress immediately. Never a path, a filename or a root name |
| `scan_settings` | `{reconcile_interval_secs, debounce_secs}` — displayed, not enforced from here |
Everything after `is_node_admin` is presentation state. It rides on the ack so a client
@@ -649,6 +663,10 @@ nodes.
* Identity keys are **per node**. There is nothing to carry between nodes, and an
operator who cracks the copy on their own disk gets a key that opens nothing
anywhere else.
+* `keypair_bundle_delete` is **reserved for `device_policy`** (`MESHBAY_DESIGN.md`
+ §3.7, open item O3): the node honours it, and no interface sends it yet. Withdrawing
+ the bundle is only safe once the account has chosen not to need it from a browser —
+ a lone button would strand the next browser that signs in.
### 7.1a Per-account blobs (MNP 3.1)
@@ -714,9 +732,9 @@ wrap_key = HKDF-SHA256(shared, salt = pk_eph, info = "meshbay:gek_wrap:v
wrapped = AES-256-GCM(wrap_key).encrypt(nonce_96, GEK, aad = pk_recipient)
```
-AES-GCM because WebCrypto has no ChaCha20-Poly1305; a `chacha20-poly1305` variant with
-`info = "meshbay:gek_wrap:v1"` exists for native clients. The recipient's public key is
-the AEAD's associated data, so a bundle cannot be re-addressed.
+AES-GCM because WebCrypto has no ChaCha20-Poly1305, and one cipher serves every
+client. The recipient's public key is the AEAD's associated data, so a bundle cannot be
+re-addressed.
`found: false` is the normal answer: per-member bundles are not stored, and the key is
produced on demand by the join path (§8). The node keeps one stored bundle of its own
@@ -1045,8 +1063,8 @@ effect immediately.
| | now - ts <= 120 s
| | rebuild A from STORED state
| | verify vs roster operator keys
- | | (file_delete also accepts the
- | | uploader's recorded key)
+ | | (file_delete also accepts any
+ | | live device of the uploader)
| | execute via ops
|<- <op>_ack {op-specific fields} --------------|
| |
@@ -1083,9 +1101,9 @@ broadcast, every connected peer in the group learns the change without reconnect
| Op | Subject | Authority | Ack | Broadcast |
|---|---|---|---|---|
-| `file_delete` | `file_id` | operator **or** the file's recorded `uploader_pk` | `file_delete_ack{file_id}` | no |
+| `file_delete` | `file_id` | operator **or** any non-revoked device of the uploading account (`uploader_id`, looked up in the roster); the recorded `uploader_pk` only when the roster cannot answer | `file_delete_ack{file_id}` | no |
| `dir_delete` | path relative to the root | operator | `dir_delete_ack{dir}` | no |
-| `invite_create` | invitee `user_id` | operator only (delegation designed, deferred) | `invite_result{code, expires_at, user_id, username}` | no — the code is shown once |
+| `invite_create` | `{user_id, username}` (canonical JSON, see below) | operator only (delegation designed, deferred) | `invite_result{code, expires_at, user_id, username}` | no — the code is shown once |
| `invite_link_create` | `link:<group_id>`, the session's group | operator only | `invite_link_result{code, invite_id, expires_at, group_id}` | no — the code is shown once |
| `invite_cancel` | `invite_id` (32 hex) | operator only | `ack{detail: "invite_cancelled", invite_id}` | no |
| `member_revoke` | `user_id` | operator | `member_revoke_ack` | no |
@@ -1093,12 +1111,13 @@ broadcast, every connected peer in the group learns the change without reconnect
| `gek_rotate` | `group_id` | operator | `gek_rotate_ack{group_id, authorized_members, note}` | no |
| `apps_enabled` | the app set | operator | `apps_enabled_ack{apps}` | yes |
| `set_scan_settings` | the interval/debounce pair | operator | `set_scan_settings_ack{...}` | yes |
-| `tmdb_config` | `custom_token=yes\|no,language=...` | operator | `tmdb_config_ack{token_customized, language}` | yes (never the token) |
+| `tmdb_config` | `{token, language}` — `token` is `null` (unchanged), `""` (clear) or `sha256:<hex>` of the token, never the token | operator | `tmdb_config_ack{token_customized, language}` | yes (never the token) |
| `tmdb_enabled` | `enabled` | operator | `tmdb_enabled_ack{enabled}` | yes |
| `tmdb_override` | `file_id=..,tmdb_id=..,media_type=..` | operator | `tmdb_override_ack{file_id, tmdb_id, media_type}` | yes |
| `tmdb_rematch` | `file_id=..` | operator | `tmdb_rematch_ack{file_id}` | yes |
| `musicbrainz_enabled` | `enabled` | operator | `musicbrainz_enabled_ack{enabled}` | yes |
-| `root_add`, `root_remove` | the root | operator | `root_add_ack` / `root_remove_ack` | no |
+| `root_add` | `{path, name, kind, writable, removable}` | operator | `root_add_ack` | no |
+| `root_remove` | the root name | operator | `root_remove_ack` | no |
| `root_update` | `<root>:rw=on\|off,rem=on\|off` | operator | `root_update_ack` | yes |
| `root_eject`, `root_plug` | the root name | operator | `root_eject_ack` / `root_plug_ack` | yes |
| `app_directories` | `<app>:<dir>,<dir>,...` | operator | `app_directories_ack{app, dirs}` | yes |
@@ -1107,7 +1126,8 @@ broadcast, every connected peer in the group learns the change without reconnect
| `search_listed` | `on\|off` | operator | `search_listed_ack{listed}` | yes |
| `transfer_limits` | `d=<n>,u=<n>` | operator | `transfer_limits_ack{limits}` | yes |
| `chat_epoch` | `group_id` | operator | `chat_epoch_ack{epoch}` | yes |
-| `group_attach`, `group_detach` | `group_id` | operator | `group_attach_ack` / `group_detach_ack` | no |
+| `group_attach` | `{name, shared_dir, writable}` | operator | `group_attach_ack` | no |
+| `group_detach` | the group name | operator | `group_detach_ack` | no |
**Upload policy is not in this table**, and that is the design: whether a member may
write is a property of each root (`root_update`), not a switch over the group. A single
@@ -1134,13 +1154,30 @@ own configuration.
A second family of operator messages is **not** signed: `node_status`, `roster_read`,
`denylist_read`, `denylist_clear`, `node_settings_set`, `node_reload`. These are gated
-by `is_node_admin()` — the authenticated session's `user_id` equals the account the
-node records as its own operator (`node_user_id`), computed from the node's own state
-and never from a hub claim. Three of them only read; the other three run through the
-same `ops` entry points as the CLI and the loopback admin API. The distinction from
-the signed table above is deliberate but worth stating plainly: a signed op proves
-possession of an operator *key*, while these prove only that the session belongs to the
-operator's *account*, which the handshake already established.
+by `_operator_device()`, which asks two things: the authenticated session's `user_id`
+is the account the node records as its own (`node_user_id`, `is_node_admin()`), **and**
+the device on this connection has proved, with `device_hello` (§9.4), a key the roster
+holds as an operator. Anything else is refused with code `not_operator`. The first
+alone would be a claim in a token the hub issued, and a hub that can name the operator
+is a hub that can be one; the second is what it cannot forge, since it holds no user
+keys. Three of them only read; the other three run through the same `ops` entry points
+as the CLI and the loopback admin API. The distinction from the signed table above:
+a signed op proves possession of an operator key for *this* operation, while these
+prove it once per connection, through the device the connection identified.
+
+**The subject covers everything the node acts on.** The signature covers `op`, the
+node, the group, the subject, the nonce and the time — nothing else of the request —
+so a value the executor uses and the subject omits is a value the operator never
+signed. Where an operation's effect is several values, the subject is canonical JSON
+of all of them (`adminop.structured_subject`, `adminSubject` in `crypto.js`: sorted
+keys, no whitespace, UTF-8), which keeps `null`, `""` and a value distinct and cannot
+be forged by a field that contains a separator. A secret is named by its SHA-256,
+because the subject is written to the audit log.
+
+**What waits for a signature is bounded.** Any authenticated member can ask for a
+challenge — the signature is checked later — so a connection holds at most 8 pending
+operations (`too_many_pending` beyond), each at most 64 KiB of subject and payload
+(`too_large`), and an expired one is dropped when the next is issued.
Rules that hold across the table:
@@ -1170,6 +1207,11 @@ one is the one that decides. A front door is allowed to differ in how it *authen
— a signature here, a run token on loopback, an operator's shell for the CLI — and never
in what it *does*.
+An operation has the doors something uses, and no more: a door nobody calls is an
+untested way in. The per-group settings the group's Settings tab changes —
+application folders, the chat folder, link previews, the Search listing, the scan
+settings — are signed MNP operations only.
+
---
## 11. Content plane
@@ -1199,7 +1241,9 @@ content-addressed: `id` is the BLAKE3 hash of the file.
| updates[], roots[]} |
|
|<- index_progress {v, group_id, scanning, | every ~2 s while scanning,
- | scanned_bytes, total_bytes} -----------| plus once on the return to idle
+ | scanned_bytes, total_bytes, | plus once on the return to idle
+ | files_done, files_total, kind, |
+ | root_pos, queued} ---------------------|
NOT sealed — see below
```
@@ -1237,7 +1281,11 @@ ejected or plugged would leave every connected client's directory table stale un
somebody reloaded the page, and the delta that tells them something changed would be the
one message unable to say what.
-`index_progress` carries counters only, never a path or filename.
+`index_progress` carries counters only, never a path, a filename or a root name:
+`kind` is what the indexer is doing (`scan`, `rescan`, `reconcile`, `watch`, or `""`
+when idle), `root_pos` is a position in the `roots` table the member already opened
+from the sealed index (`-1` when none), and `queued` is the number of roots waiting
+their turn, not their names.
### 11.1a The sealed envelope
@@ -1267,9 +1315,9 @@ ciphertext. That is the same statement the index makes, one step stronger.
`salt = <none>` is Python's `salt=None` and WebCrypto's `salt: new Uint8Array(0)`;
RFC 5869 extracts with a zero key either way. The subkeys are purpose-separated
-rather than borrowed from a file's key space — `GroupIndex.serialize()` reuses
-`chunk_key_aes` with a pseudo-file ("the index as chunk 0 of a virtual index file"),
-which is a hack this deliberately does not repeat.
+rather than borrowed from a file's key space — reusing `chunk_key_aes` with a
+pseudo-file ("the index as chunk 0 of a virtual index file") is a hack this
+deliberately does not repeat.
**What stays in clear, and why each one has to:**
@@ -1314,10 +1362,6 @@ purpose and a fresh 96-bit random nonce per message: at one message per 48 KiB c
reaches 2⁻³², and `gek_rotate` exists. Deriving the nonce from the payload instead
would be worse, not better — two chunks of identical bytes are ordinary in a file.
-`GroupIndex.serialize()` is not a candidate for reuse here: it compresses with zstd,
-which no browser can decompress (`DecompressionStream` offers gzip and deflate only),
-so reusing it would mean shipping a WASM decoder to every client for no gain.
-
**Failure is fatal, never degraded** (I8). A client that cannot open an index message
ends the session naming the message type; it never reports an empty index, because "the
group has no files" is a state a real group can be in.
@@ -1376,8 +1420,8 @@ an absent setting as no limit would leave the node-wide cap as the only control,
is the situation leases exist to end. The per-group value is a signed operator
operation (`transfer_limits`, §10.4, bounded to 1–32; zero is refused, because a member
who may not transfer at all is a member the operator revokes). The node-wide values are
-daemon settings. A member's own caps ride on the handshake ack so the interface can say
-"2 of your 2 slots are busy" rather than draw a spinner that explains nothing.
+daemon settings. A member's own cap rides on every `transfer_state`, so the interface can
+say "2 of your 2 slots are busy" rather than draw a spinner that explains nothing.
**Queueing.** One FIFO per kind. `_pump` walks it in arrival order and **skips** a
member who is at their own cap rather than stopping at them — granting strictly in
@@ -1386,9 +1430,9 @@ told how many are `ahead` of it. Beyond 32 queued per member the answer is
`too_many_queued`, because an unbounded queue is how a node runs out of memory politely.
`used` and `cap` on `transfer_state` are this member's own count and this member's own
-limit **in this group** — the same value `_has_room` enforces and the same one the
-handshake ack announces. Three readings of one number, and an interface that draws a
-different one from the node's is an interface that offers a slot the node will queue.
+limit **in this group** — the same value `_has_room` enforces. The interface reads it
+from here and from nowhere else: one number with two sources is an interface that can
+offer a slot the node will queue.
**Reclaim.** The session teardown is the primary path and it is immediate. A sweeper
runs every 15 s for whatever the teardown cannot see, and tells two failures apart:
@@ -1419,6 +1463,19 @@ numbers), `bad_kind`, `too_many_queued`, and `not_your_transfer` — the last fo
or closing a `tr` another connection holds, which would otherwise be a denial of
service one random id away.
+**A lease is what the node granted, not what the client called it.** `tr` is drawn by
+the client, so on `file_req` and `file_upload` it is a claim, resolved against the
+node's own record for this connection:
+
+| `tr` names | `file_req` | `file_upload` |
+|---|---|---|
+| a lease of this connection, **granted** | served as a leased transfer; marks the lease alive | accepted; marks the lease alive |
+| a lease of this connection, still **queued** | refused, `lease_not_granted` — reading while queued is the cap not applying | refused, `lease_not_granted` |
+| nothing this connection holds, or no `tr` | a leaseless read, counted against the ceiling below | accepted, bounded by the upload protections (§11.4) — this is what a reconnect looks like, with the old connection's leases gone and the client re-opening them |
+
+Read as a bare presence check, the field would let any non-empty string skip the
+leaseless ceiling and every cap behind it.
+
#### Reads that carry no lease
Browsing a group is **never** subject to a transfer slot: not the poster grid, not the
@@ -1465,15 +1522,16 @@ as one.
```
C N
|-- file_req {v, file_id, chunk_index, [tr]} ------->|
- | | `tr` present: mark that lease
- | | alive (§11.2)
+ | | `tr` of a granted lease: mark
+ | | it alive; of a queued one:
+ | | `lease_not_granted` (§11.2)
| | index lookup; if the id is
| | not a file, try the media
| | cache (thumbnail/poster/
| | cover/transcode), sliced the
| | same way — never leased
- | | `tr` absent: admit against the
- | | leaseless ceiling, else
+ | | no granted lease: admit against
+ | | the leaseless ceiling, else
| | `transfer_required`
| | backpressure: wait while
| | bufferedAmount > 2 MiB
@@ -1494,8 +1552,9 @@ nonce = 12 random bytes
ct = AES-256-GCM(chunk_key).encrypt(nonce, plaintext) no AAD
```
-* The `:aes` suffix keeps AES keys distinct from the ChaCha20 variant
- (`chunk_key`/`encrypt_chunk`, `info` without the suffix) derived from the same GEK.
+* The `:aes` suffix is part of every chunk key: it once kept these distinct from a
+ ChaCha20 variant derived from the same GEK, which no longer exists, and it stays
+ because removing it would change every key.
* `nonce` and `ct` are msgpack **binary**, not base64. Every message that carries
content carries it this way, and none carries it outside an AEAD.
* The key is a pure function of (GEK, file hash, index), so chunks are cacheable,
@@ -1515,6 +1574,11 @@ ct = AES-256-GCM(chunk_key).encrypt(nonce, plaintext) no AAD
* **A chunk that is not this shape aborts the download**, with an error. There is no
fallback that decodes it some other way: a client that guesses at a chunk it does not
recognise writes its guess into the file the person is saving.
+* **In a public group, a file the hub's content blocklist names is not served.** It is
+ left out of `index_sync` and `index_delta`, and `file_req` for it or for its
+ thumbnail, `stream_req`, `subtitle_req` and `audio_transcode_req` are refused with
+ `content_blocked`. The node syncs the list from the hub and applies pushed changes
+ (`MESHBAY_DESIGN.md` §7.5); private groups are never affected.
### 11.4 Upload
@@ -1852,6 +1916,8 @@ duplicates is already stored and there is nothing for anyone to retry.
| the connection has identified its device (§9.4) | the `device` field is what receivers verify against |
| `device` == this connection's pinned device | **a device may only send as itself.** A member free to name another member's key could *be* that member to everyone, and the signature would verify |
| `sender_id` is taken from the authenticated session | never read from the message; it never was |
+| ciphertext at most 64 KiB (`chat_too_large`) | a message is a row on the operator's disk that nothing expires, a relayed copy for every connected member and a notification for every member |
+| at most 60 messages per 60 s per account per group (`chat_rate_limited`) | keyed by account, not connection, so a second tab does not buy a second budget. There is no node-wide chat ceiling: it would let a busy group silence a quiet one |
`format` is a *storage* state as well as a wire one, so the store can distinguish rows
the wire would not accept. Nothing changes what is accepted from a peer: the sealed
@@ -1994,7 +2060,6 @@ it back (§3.5).
| `stream_more` / `stream_stop` | C→N | auth | grant credit / abandon the stream |
| `chat_msg` | C→N, N⇒C | auth | send and fan out a message |
| `chat_hist` / `chat_hist_resp` | C→N / N→C | auth | paged history |
-| `chat_attach` | — | auth | attachment metadata (declared, unused on the wire) |
| `link_preview_req` / `_resp` | C→N / N→C | auth | OpenGraph unfurl |
| `media_meta_req` / `_resp` | C→N / N→C | auth | TMDB metadata for one file |
| `season_meta_req` / `_resp` | C→N / N→C | auth | per-season TMDB fields |
@@ -2003,6 +2068,9 @@ it back (§3.5).
| `audio_transcode_req` / `_resp` | C→N / N→C | auth | browser-playable copy of a WMA/MPC file |
| `subtitle_req` / `_resp` | C→N / N→C | auth | one embedded subtitle track as WebVTT, by cache hash |
| `ping` / `pong` | C→N / N→C | auth | liveness on an open channel |
+| `admin_challenge` | N→C | auth | the transcript fields of a signed op, to rebuild and sign (§10.2) |
+| `admin_response` | C→N | auth | the operator's signature over that transcript |
+| `client_diag` | C→N | auth | the video player's own view of a stream, written to the node's log beside its own (a stream event at INFO, the periodic state at DEBUG); the node acts on none of it and sends no reply |
| `member_revoke` / `_ack` | C→N / N→C | signed | stop serving the key to someone |
| `member_unpin` / `_ack` | C→N / N→C | signed | forget a pinned identity |
| `transfer_limits` / `_ack` | C→N / N⇒C | signed | per-member transfer caps for this group |
@@ -2029,7 +2097,6 @@ it back (§3.5).
| `denylist_clear` / `_ack` | C→N / N→C | auth (operator) | remove entries |
| `node_settings_set` / `_ack` | C→N / N→C | auth (operator) | change daemon settings |
| `node_reload` / `_ack` | C→N / N→C | auth (operator) | re-read `node.toml` |
-| `ephemeral_stream` | — | — | reserved, mobile live push |
| `error` | N→C | any | refusal, with `detail` and optionally `code`, `req_id`, and the `upload_id` / `tr` / `file_id` it is about |
| `ack` | N→C | auth | generic acknowledgement (chat, keypair bundle store) |
@@ -2052,7 +2119,7 @@ speak, and every message above is available on it.
**QUIC is in development** (§5.2): a partial message set, no client, and not a shipped
feature. Nothing about it is a compatibility commitment yet.
-Two rules hold across transports, and both are about there being exactly one of each
+One rule holds across transports, and it is about there being exactly one of each
message:
* **One encoder per message type, shared by every transport.** `file_chunk` comes from
@@ -2061,21 +2128,25 @@ message:
its own. Two encoders for one type is a type free to drift, with a name that no longer
says which shape will arrive — and, when one of them is a sealed envelope, a second
construction site that goes on sending cleartext.
-* **`GroupIndex.serialize()` / `deserialize()` describes no MNP message.** It is a
- signed, compressed, encrypted at-rest and interchange format, and reading it as a wire
- contract is a mistake worth naming: the sealed envelope of §11.1a is what index
- messages travel under, and it is deliberately not this, because zstd decompresses in
- no browser.
---
## 13. Versioning and compatibility
-MNP versions independently of the package version. Current: **`3.4`**; oldest peer
-accepted: **`3.0`** — 3.1, 3.2, 3.3 and 3.4 are all additive, so the floor does not move
-with them. 3.4 adds the challenge signature (section 6.5) and invitation links
-(section 8.6): `invite_link_create`, `invite_link_result`, `invite_cancel`, which an older
-node answers as unknown messages.
+MNP versions independently of the package version. Current: **`5.0`**; oldest peer
+accepted: **`4.0`**.
+
+4.0 is the floor: a member presents a short-lived node-audience token bound to one node
+(§6.3) instead of its hub session token, a change to what a peer must *present*, so a
+pre-4.0 client is refused at the handshake. It carries everything 3.x added — the
+challenge signature (§6.5), invitation links (§8.6), audio-track and subtitle
+selection, per-account blobs.
+
+5.0 is a MAJOR confined to four signed operations — `root_add`, `group_attach`,
+`invite_create`, `tmdb_config` — whose subjects now name every value the node acts on
+(§10.4). A peer across the break refuses to sign the other side's subject, so those
+four fail with a refusal and everything else works; no node accepts the old subjects,
+so nothing is left unsigned on either side. That is why the floor did not move.
The two numbers are separate on purpose. `MNP_VERSION` says what this build speaks;
`MNP_MIN_SUPPORTED` says what it will talk to, and moving the second is a decision about
@@ -2157,7 +2228,9 @@ walks through the gate meant to stop it.
| No MitM on the signaling path | both DTLS fingerprints (or the QUIC certificate hash) inside every transcript; empty binding is a refusal |
| The hub cannot read content | the group key never reaches it; chunk keys and chat epoch keys derive from or are wrapped under it |
| The hub cannot substitute a key at invite | the node wraps for a key its owner presented and signed, bound to an identity by a code the hub never sees |
-| The hub cannot administer a node | privileged ops need an Ed25519 signature from a roster-pinned operator key |
+| The hub cannot administer a node | privileged ops need an Ed25519 signature from a roster-pinned operator key; the unsigned operator controls need a device that proved such a key (§10.4) |
+| The token a member hands a node opens nothing at the hub | the handshake takes only `aud = MNP_AUD` tokens, and the hub API only its own audience (§6.3) |
+| A token captured by one node's operator is useless at another node | the MNP token names the node it was minted for, and a node refuses one naming a different key (`wrong_node`) — before the pre-proof window can serve anything |
| The hub cannot impersonate a device owner | device linking is countersigned by a key the node pinned; the hub stores no user keys |
| A revoked token cannot connect | denylist on user, `jti` and group, persisted across restarts |
| A signature cannot be repurposed | domain-separated, length-prefixed transcripts naming op, subject, node, group, nonce and time |
@@ -2183,9 +2256,10 @@ walks through the gate meant to stop it.
attacks apart: the pairing code defeats a hub that *lies in its directory* — silent,
undetectable, per-request — not one that *rewrites the client*, which is an artifact
that can be inspected and compared.
-* **An invitation link's code is a bearer code** (§8.6). The node admits whoever brings
- it first; what restricts who can bring it is the hub, which lets only the addressed
- account reach the node — a rule an active hub does not have to keep. The challenge
+* **An invitation link is a bearer secret on both halves** (§8.6). The hub admits the
+ first account that redeems its ticket, and the node admits whoever brings the code
+ first; the link is bound to no address, so whoever holds it first joins. What bounds
+ it is that it works once, for seven days, and can be cancelled. The challenge
signature (§6.5) keeps the code from reaching any node but the issuer; it does not
keep it from a hub the inviter asked to mail it, which then holds it.
* **The pre-proof window is a disclosure surface.** A hub that forges a JWT can fetch a
@@ -2281,8 +2355,10 @@ LP(x) = uint32be(len(x)) || x every field, no exceptions
| Constant | Value | Source |
|---|---|---|
-| `MNP_VERSION` | `3.4` | `meshbay_common/__init__.py` |
-| `MNP_MIN_SUPPORTED` | `3.0` | `handshake.py` |
+| `MNP_VERSION` | `5.0` | `meshbay_common/__init__.py` |
+| `MNP_MIN_SUPPORTED` | `4.0` | `handshake.py` |
+| `MNP_AUD` / `HUB_API_AUD` | `meshbay:mnp` / `meshbay:hub-api` | `tokens.py` |
+| MNP token lifetime | 900 s | `meshbay-hub/auth.py` (`issue_mnp_token`) |
| `NONCE_LEN` | 32 bytes (both handshake nonces) | `handshake.py` |
| `ADMIN_CHALLENGE_TTL` | 120 s | `adminop.py` |
| `JOIN_TTL`, `DEVICE_TTL` | 120 s | `join.py`, `device.py` |
@@ -2317,10 +2393,11 @@ LP(x) = uint32be(len(x)) || x every field, no exceptions
| `MAX_CONCURRENT_TRANSCODES` | 8 | `webrtc/apps/streaming.py` |
| Link preview rate | 15/conn, 60/node per 60 s; cache 1 h × 256 | `webrtc/chat.py` |
| ICE gathering deadline | 4 s | `transport.js` |
-| Signaling: max SDP, pending per user, rate | 16 KiB, 3, 30/min, 15 s answer timeout | `api/signaling.py` |
+| Signaling: max SDP, pending per account, per-account-per-node budget, per-address rate | 16 KiB, 32, burst 120 refilled at 2/s, 600/min per node, 15 s answer timeout | `api/signaling.py` |
+| `MAX_CHAT_CIPHERTEXT` / chat rate | 64 KiB / 60 per 60 s per account per group | `webrtc/chat.py` |
| GEK | 256-bit, node CSPRNG | `crypto.py` |
| Chat epoch key | 256-bit, node CSPRNG, one per group per epoch | `chatbox.py` |
-| Chunk cipher | AES-256-GCM, 96-bit nonce (ChaCha20-Poly1305 variant for native) | `webcrypto.py`, `crypto.py` |
+| Chunk cipher | AES-256-GCM, 96-bit nonce | `webcrypto.py` |
| GEK wrap | X25519 + HKDF-SHA256 + AES-256-GCM, AAD = recipient public key | `crypto.py` |
| Sealed message envelope | HKDF-SHA256 subkey per purpose, AES-256-GCM, 96-bit random nonce | `groupbox.py`, `static/crypto.js` |
| Chat envelope | HKDF-SHA256 subkey per device per epoch, AES-256-GCM, 96-bit random nonce, Ed25519 over the ciphertext | `chatbox.py` |
@@ -2340,8 +2417,9 @@ meshbay-common/ protocol.py message types, chunk and upload codecs
adminop.py the admin transcript and the operation catalogue
join.py join transcript and pairing codes
device.py device request / add / hello transcripts
- crypto.py GEK, chunk keys, ECIES wrap, BLAKE3 ids
- webcrypto.py the AES variants the browser can also compute
+ crypto.py GEK, ECIES wrap, keystore, BLAKE3 ids
+ webcrypto.py the content cipher: per-chunk AES-GCM keys
+ tokens.py the two token audiences (node vs hub API)
meshbay-node/ transport/webrtc_server.py the reference implementation of MNP,
transport/webrtc/ assembled from these modules
diff --git a/docs/playlists.md b/docs/playlists.md
index f4a2c28..c41c71f 100644
--- a/docs/playlists.md
+++ b/docs/playlists.md
@@ -154,9 +154,8 @@ result with no hub change at all.
### 3.2 Not node-to-node
-Refused, and not on cost grounds. Nodes do not know each other, share no
-authenticated channel, and `replication.py` is legacy public-content code that
-has nothing to do with this. Beyond the protocol that would have to be
+Refused, and not on cost grounds. Nodes do not know each other and share no
+authenticated channel. Beyond the protocol that would have to be
invented, it leaks the thing this architecture is most careful about: node A
would learn that this account also uses node B — that two unrelated operators
host the same person. Per-node identity exists precisely so that this