diff options
Diffstat (limited to 'docs/MESHBAY_DESIGN.md')
| -rw-r--r-- | docs/MESHBAY_DESIGN.md | 224 |
1 files changed, 152 insertions, 72 deletions
diff --git a/docs/MESHBAY_DESIGN.md b/docs/MESHBAY_DESIGN.md index 78376a3..2efedca 100644 --- a/docs/MESHBAY_DESIGN.md +++ b/docs/MESHBAY_DESIGN.md @@ -16,7 +16,7 @@ > them — it names the invariant that holds today, not the incident that produced > it. §13 is the register of those labels. > -> Wire versions at the time of writing: **MNP 4.0** (oldest peer accepted 4.0), +> Wire versions at the time of writing: **MNP 5.0** (oldest peer accepted 4.0), > **MHP 0.1**, packages **0.16.0**. The normative source for the wire format is > `MESHBAY_NODE_PROTOCOL.md`; this document states the design the protocol > serves, not its byte layout. @@ -99,7 +99,7 @@ opens them. │ └─────────┘ MHP 0.1 │ signalling (SDP/ICE, <1 KB), presence, revocation push MHP │ - ┌────┴────┐ MNP 3.0 ┌──────────┐ + ┌────┴────┐ MNP 5.0 ┌──────────┐ │ node │◄──────── WebRTC DataChannel / QUIC ──────────►│ client │ └─────────┘ index, file chunks, streams, chat, admin └──────────┘ holds the files browser SPA or desktop @@ -128,8 +128,9 @@ else about content. This is decision **E9**, and it is the rule any new feature is measured against. A feature that wants a row on the hub about a group's content is a feature that has misunderstood the model. It has been re-verified at each content-model change: -`SwarmSource` carries a content hash, a node id and an endpoint — **no paths, no -filenames** — and private groups register nothing at all (**H7**). +**no node registers a content hash with the hub**, for any group (**H7**); the only +hashes the hub holds are those of public content somebody reported (§7.5) — **no +paths, no filenames**. --- @@ -293,7 +294,7 @@ comes from the hash binding: |---|---| | Request TTL | 1 h, `[node] device_request_ttl_minutes` | | Devices per account per node | 5 | -| Attempts per connection | 5, then a node-wide lockout | +| Attempts per connection | 5, audited on exhaustion | | Filing a request | requires the account to have at least one pinned identity already | `member unpin <user>` removes **every** device of an account, and is the only @@ -489,9 +490,10 @@ The Members tab can ask the hub to mail the code to the invitee's address on fil can join in the invitee's place. It is offered because a code that arrives on its own is worth more to most groups than the property, and it is stated rather than hidden: the box reads *"Send the invitation by e-mail (may land in spam)"*, is -checked by default, and is **remembered per account** (the `invite_email` -preference), so an operator who unticks it once is not asked to again. Unticked, -the hub is never called and the table above holds exactly. The CLI mails nothing. +**unticked by default** — giving the hub the code is something the inviter opts into, +never something done unasked — and is **remembered per account** (the `invite_email` +preference), so an operator who ticks it once is not asked to again. Unticked, the +hub is never called and the table above holds exactly. The CLI mails nothing. The same box sits under the link form, sharing the same preference. Ticked, and with an address typed, the hub mails the link to that address — and @@ -519,8 +521,7 @@ group key.** That is a property of open joining, not a defect of this design. Content in such a group is protected from the network and from non-members, and from nobody else. -Note the axis. **`visibility`** (public/private) controls discoverability and swarm -hash registration (**H7**). **`join_policy`** (open/request/invite) controls +Note the axis. **`visibility`** (public/private) controls discoverability. **`join_policy`** (open/request/invite) controls admission. Only the second decides whether a code is required: a public group with `join_policy = "invite"` keeps the code, because being findable is not being open. @@ -549,8 +550,10 @@ rendered as a grouped mnemonic. `recovery_key = HKDF-SHA256(R, info = "meshbay:recovery:v1:" + username)` — HKDF and not Argon2, because `R` has 256 bits and there is nothing to brute-force. Every time an identity bundle is written to a node, a **second copy** is written beside it wrapped under `recovery_key` -(`bundle_enc_recovery`, additive on the wire). `R` is a pass-through: offered in -the registration email by default, never written to any database, never logged. +(`bundle_enc_recovery`, additive on the wire). `R` is a pass-through: shown once +on screen at registration and mailed with the verification code only if the person +ticks the box for it (unticked by default), never written to any database, never +logged. The reset endpoints are built to leak nothing. `POST /v1/users/password/reset-request` requires **the username and the email on file as a pair**, checked against a blind @@ -615,7 +618,8 @@ control. It closes for a native device unconditionally, because that device's ke is in no bundle anywhere. It closes for an *account* only when no browser needs a bundle on that node — which needs `device_policy {allow_bundle: false}`, **signed by a pinned key** so the decision is the user's and never the hub's (open item -O3). +O3). Withdrawing a bundle already exists on the wire (`keypair_bundle_delete`) and +is not offered in the interface: it belongs with that decision, not before it. --- @@ -635,17 +639,17 @@ Every private key lives in an encrypted keystore on the machine that owns it. Th hub never sees one. **Domain separation is consistent and mandatory.** Every derivation uses a -distinct `info` string, and the AES variant adds an `:aes` suffix so two ciphers -can never derive the same key from one group key. This is a small detail that -prevents cross-protocol key reuse, and it is checked rather than assumed. +distinct `info` string, so no two purposes can derive the same key from one group +key. The chunk and wrap strings end in `:aes`, left from a second cipher that no +longer exists; it stays because it is part of every key already derived. ### 4.2 Group key wrapping (ECIES) ``` wrap: sk_eph, pk_eph = X25519.generate() # fresh per bundle shared = X25519(sk_eph, pk_recipient) - wrap_key = HKDF(shared, salt=pk_eph, info="meshbay:gek_wrap:v1", len=32) - wrapped = AEAD(wrap_key).encrypt(nonce, gek, aad=pk_recipient) + wrap_key = HKDF(shared, salt=pk_eph, info="meshbay:gek_wrap:v1:aes", len=32) + wrapped = AES-256-GCM(wrap_key).encrypt(nonce, gek, aad=pk_recipient) bundle = pk_eph ‖ nonce ‖ wrapped unwrap: shared = X25519(sk_recipient, pk_eph) # same derivation @@ -675,20 +679,19 @@ time. This avoids double storage and makes key rotation feasible without re-encrypting terabytes. ``` -disk (plaintext) → compress → per-chunk AEAD under a group-derived key → transport → client +disk (plaintext) → per-chunk AES-256-GCM under a group-derived key → transport → client ``` - Chunk size 1 MB: amortises AEAD overhead and enables seeking, because each chunk is independently decryptable. -- `chunk_key = HKDF(GEK, salt=None, info="file:" ‖ blake3(file) ‖ ":chunk:" ‖ index)`. +- `chunk_key = HKDF(GEK, salt=None, info="file:" ‖ blake3(file) ‖ ":chunk:" ‖ index ‖ ":aes")`. The salt is omitted deliberately: the group key is CSPRNG output and already uniform, so the file and chunk context belongs in `info`, which is the correct HKDF usage (**M5**, first review). - **Chunk authentication is the AEAD tag**, not a per-chunk signature. The tag authenticates the ciphertext under a key only members hold, which is what the signature was for. -- Compression precedes encryption, because compression is ineffective on - ciphertext. +- **Chunks are not compressed**: a chunk is encrypted and sent as it was read. - Upload chunk size is 48 KB, which is what fits the SCTP limit after msgpack overhead. @@ -929,7 +932,7 @@ implementations of one security check is **C6** waiting to happen. ``` client → node handshake {token, group_id, nonce_c, v, v_min} -node authorize_token() JWT · scope · denylist · group_id · membership · hosting +node authorize_token() JWT · aud · scope · node · denylist · group_id · membership · hosting node → client handshake_challenge {nonce_s, node_pk, sig} sig: Ed25519 over the challenge (3.4) ── pre-proof window: bundle fetch, join ── client → node handshake_response {proof} @@ -956,7 +959,9 @@ buys, per the convention at the top: a client that knows which node it means to reach can refuse to send a code anywhere else — against a hijacked signaling path and against a second host of the same group. It proves *a* key, not the *right* one: it helps only a client that already knows which key to expect. A wrong -signature is refused; an absent one is an older node, discovered from its answer. +signature is refused, and so is an absent one: every node the floor admits signs +whenever it has a channel binding, and one without a binding could not complete the +proof anyway. **Channel binding is mandatory and an absent one is refused** — never degraded to nonce-only, which would silently drop MitM detection: @@ -1000,7 +1005,9 @@ for `hubFetch` and signaling, and a short-lived **MNP token** (`aud = MNP_AUD`, from `POST /v1/nodes/mnp-token`) that carries the member's `sub`, `groups` and `jti` and is the only thing presented in the handshake. The node binds `MNP_AUD` when it decodes, so a session token is refused here; the hub API binds its own -audience, so an MNP token captured by an operator is refused there. It is +audience and requires `exp`, `sub` and a known `scope`, so an MNP token captured +by an operator is refused there, and so is any other hub-signed token that is not a +session (a revocation broadcast, an MHP token). It is checked once, before the proof, so its short life never interrupts a transfer or a film already playing — a reconnect fetches a fresh one. `meshbay_common/tokens.py` holds the two audience strings, shared by the hub that issues and the node that @@ -1031,10 +1038,11 @@ requester and it therefore grants nothing across accounts. `gek_required: false` bypass (**NS8**). **Refusals carry a code**, not only a sentence, because a client can act on a code. -`not_a_member` in particular is usually a token issued before the person was added -to the group — `groups` is baked in at sign-in and the hub pushes no updates — so -the client refreshes once and retries rather than telling someone who was invited a -minute ago that they are not a member. +`not_a_member` means the hub did not count the account a member when it minted the +token; since the MNP token is minted per connection from the membership the hub holds +then, a stale `groups` claim is no longer the usual cause. The client still refreshes +once and retries on that code before telling someone who was invited a minute ago +that they are not a member. ### 5.3 Correlation and liveness @@ -1063,7 +1071,15 @@ TTL 120 s. **The client reconstructs the transcript from announced fields and refuses to sign if the operation or subject is not what the user asked for** (**H5**) — a challenge of opaque random bytes signed blind is an unbound signing oracle. The transcript's subject names the *outcome*, not the operation: what the -operator is shown before signing has to be what happens. +operator is shown before signing has to be what happens. **Every value the node acts +on is in the subject**: the signature covers nothing else of the request, so an +operation whose effect is several values signs all of them as canonical JSON — a +root's path *and* whether every member may write there, a group's name *and* the +directory it exposes — and a secret by its SHA-256, since the subject is audited. + +What waits for a signature is bounded: anyone authenticated can ask for a challenge, +so a connection holds at most eight pending, each at most 64 KiB, expired ones +dropped (**AV31**). Verification is against `roster.operator_pks()`, rebuilt from node state, **never** from anything in the response. @@ -1271,19 +1287,21 @@ new client, and the hub, node and SPA deploy together. with each MAJOR, so `check_version` refuses at the handshake any peer that cannot meet one: an upload is sealed or it is not sent; a transfer has a real lease or it does not run; there is one app-directories op and no wrappers behind it. The client -records the version its peer declared, for diagnostics, and **branches on none of -it**. +checks the version its peer declared and **branches on none of it**. > A capability flag on a peer whose floor already guarantees the capability is a > branch that can only ever take one path — until somebody lowers the floor, at > which point it silently takes the other. **A field kept "just in case" is how > the branches come back.** -**The floor is not the current version, and MINOR additions are why.** It is -`MNP_MIN_SUPPORTED` in `handshake.py`, it equals the last MAJOR, and 3.1, 3.2, -3.3 and 3.4 have all been added above it without moving it. So a peer can be reachable and -still not do something the current version can, and the client has to cope with -that — **by reading the peer's own answer, never by comparing version numbers**. +**The floor is not necessarily the current version, and what is added above it is why.** +It is `MNP_MIN_SUPPORTED` in `handshake.py`, and it is the last MAJOR that had to +refuse at the handshake: 3.1–3.4 were added above the 3.0 floor without moving it, +and 5.0, a MAJOR confined to four signed operations that a peer across the break +refuses to sign, sits above the 4.0 floor. So a +peer can be reachable and still not do something the current version can, and the +client has to cope with that — **by reading the peer's own answer, never by comparing +version numbers**. 3.2's audio tracks are the worked example: the node lists them in `stream_init`, the client draws its selector from that list, and a node that sends no list gets no selector. 3.3's subtitles repeat it exactly, and add the case where the list is @@ -1813,8 +1831,10 @@ the next node on a `not_hosted` refusal (`MESHBAY_NODE_PROTOCOL.md` §6.3). The node authenticates to the hub with an Ed25519 signature over a domain-separated timestamped message — **no password and no auth key on a node** — and receives a -`scope: "node"` token that is refused for group management. The operator manages -groups from a client (**NS7**). +`scope: "node"` token that is refused for group management and on every admin and +moderator route, even when the account behind it holds a hub role: what a node may do +is its operator's roster pin, and a hub role is a person's, exercised from a client. +The operator manages groups from a client (**NS7**). Signaling is rate-limited, SDP-size bounded, capped per user, and **the caller must share an active group with the target node**. Otherwise any authenticated user @@ -1930,8 +1950,9 @@ listed with what they check. The hub publishes no API description — no `/docs` `/redoc` or `/openapi.json` — in the code, not in a proxy rule, so a packaged install behind any proxy publishes none either. -The mail bounds (`mail.*`) and the sign-in lockout (`login.max_failures`, -default 4, and `login.lockout_minutes`, default 60 — §7.7) live in the same table +The mail bounds (`mail.*`), the report policy (`reports.*`, §7.5) and the sign-in +lockout (`login.max_failures`, default 4, and `login.lockout_minutes`, default 60 — +§7.7) live in the same table for the same reason: they are what an operator changes while the hub is serving, from the panel, without a restart. Each value is clamped to published bounds, and `max_failures = 0` turns the lockout off. @@ -1950,17 +1971,71 @@ The client shows the real state, not a blanket one. **Revocation is honoured by nodes** and the denylist survives a restart (**H4**); signaling refuses a group that is not active. +**Revocation has one door, and it broadcasts.** Only an administrator revokes, and +only through `POST /v1/admin/revoke`, which signs the revocation and pushes it to every +connected node — the same signed broadcast an administrator's account deletion sends +(§7.7). The user and group PATCH handlers refuse `revoked` outright, because a +status written there reached no node and behaved as a suspension while claiming to be +a revocation. Moving a group or an account *out* of `revoked` is an administrator's +call too, and it changes the hub row only: the nodes keep enforcing the revocation +they received, so the hub and the nodes then disagree until each operator clears it +(`denylist clear`). That is why the table says "no". + **Moderator is not administrator.** The user-patch handler is split by field: a moderator may act on the fields moderation needs and may not write `role`. -**Content reporting requires authentication, distinct reporters and a rate limit, -and is refused when public groups are off.** An unauthenticated endpoint that -blocklists a content hash after two reports is a network-wide censorship and DoS -primitive for anyone who learns a public file's id. +**A file is reported by a member of the public group it was seen in, and an +administrator decides.** An unauthenticated endpoint that blocklists a content hash +after two reports is a network-wide censorship and DoS primitive for anyone who +learns a public file's id, so every bound here answers what a report costs somebody +else — a file taken out of a group everyone else uses, and an administrator's time: + +- a **person's** account — a node's token is refused — that has existed for + `reports.min_account_age_hours` (24 by default); +- **active membership of the public group** named in the report, answered with one + refusal whatever the reason, so the endpoint says nothing about which groups exist + or who is in them. The hub cannot check that the file is in that group — it holds + no index, by design — only that the reporter could have seen it there; +- a **daily allowance per account** (`reports.daily_per_account`, 20) besides the + per-address rate limit, since an address is one of thousands a subscriber holds; + one report per account per hash; a closed list of reasons and a bounded detail; +- once `reports.review_threshold` distinct accounts (3) have reported a hash, it is + **queued for review** and the administrators are notified. An administrator blocks + it — which pushes it to the nodes — or dismisses it, and a dismissed hash is not + reopened by more reports. The reporter is told the report was recorded and never + how close the file is to review. + +`reports.auto_block` blocks at the threshold without a review. It is off by +default and stays an instance's explicit choice, because it makes a handful of +accounts made for the purpose enough to take a file down. The whole flow is refused +while public groups are switched off. The interface offers **Report** on a file's +menu in a public group only. -The exact-hash CSAM check is **structural, not yet functional** — production -databases are perceptual — and is stated as such so it is not relied on -operationally. +**The content blocklist is applied by the nodes that host a public group, in their +public groups only.** A node holds the list (`blocklist.py`, persisted beside the +denylist so a restart while the hub is unreachable does not serve again what had +stopped being served), fetches the whole of it on every connection to the hub — +`GET /v1/blocklist`, paged by hash, answered to a node's own token only — and +receives each addition and removal pushed on its hub socket (`blocklist_update`), +sent only to nodes registered for a public group. In a public group a blocked file +leaves the index members are sent, and a request for it, its thumbnail, a stream of +it, a subtitle track or an audio conversion of it is refused (`content_blocked`); +the members connected when the list changes are resent the index. Nothing is +deleted: the file is on the operator's disk, and what they keep is theirs. + +Stated per the convention at the top: + +- **Private groups are untouched**, by construction: no node sends the hub a + content hash (**H7**), so nothing in a private group can be on the list. +- **Which groups are public is the node's own configuration.** The hub and the + node are given the same value when a group is created and the hub never changes + it; an operator who edits `node.toml` to call a hub-listed group private takes + it out of the list's reach. +- **It is an exact match on the content id.** A file changed by one byte is + another id, and a file past the partial-hash threshold is identified by a sample + of its bytes (§6.3). It is a moderation tool, not a guarantee. +- **The QUIC transport does not apply it** — it is in development and serves no + client (§5.1, §15.3). ### 7.6 Federation (MHP) @@ -2002,7 +2077,7 @@ radius is proof of the passphrase. An admin can delete one too. The row is **tombstoned rather than dropped**: username released, email and password hash cleared, node linking key dropped, memberships, notifications, refresh tokens, -node registrations, device keys and public-swarm sources removed, active tokens +node registrations and device keys removed, active tokens refused at once by a status check rather than left to expire. Device keys go because the desktop client keeps its half: left on the tombstone, the key would refuse that installation to the next account created from it. @@ -2067,8 +2142,8 @@ The rules that make this safe: attempts than the limit. A request that checked no passphrase gives its attempt back. - **Every path that checks the passphrase counts on the same row**: sign-in, - passphrase change and account deletion. A right passphrase clears it; failures - older than the window age out. + passphrase change, changing the e-mail address on file and account deletion. A + right passphrase clears it; failures older than the window age out. - **A lockout refuses passphrase sign-in and nothing else.** Open sessions, token renewal and device sign-in continue, and a reset code sent to the address on file clears it — so a stranger who locks a public username costs its owner at @@ -2462,7 +2537,7 @@ everything registered — a node that predates an application hides nothing. **An application's directories are the same shape one level down**: one generic signed op (`app_directories`) keyed by the application's own registry name, stored -under `<key>_directories`, one MNP message, one loopback route. Adding an +under `<key>_directories`, one MNP message. Adding an application adds **no function, no message type and no route** — which is what "plug-in architecture" has to mean to be worth the phrase. @@ -2908,6 +2983,15 @@ requirements, not compatibility notes. | **Watcher reliability** | Change notification drops events under load on Windows, and inotify is unreliable on a FUSE mount. **Periodic reconciliation is mandatory on both platforms** | | **No symlinks, no POSIX permissions** | Simplifications: nothing to defend against, and the node runs as the user anyway | +**The client makes a name writable when it saves, and says so.** A node serves +the name its disk gave the file and never rewrites it — that is the string that +opens it — so a single file, a zip's entries and the zip's own name are passed +through `portable-name.js` at the moment of saving (reserved characters become +`_`, a trailing dot or space goes, a reserved stem gains `_`), and the transfer's +row names the original when it changed. The rule is `paths.sanitize_for_download`, +and the two are held byte-identical by a parity test. Two different names can +still become one — a zip keeps both entries under it. + **Case folding is for comparisons the code makes itself** — index identity, collision reporting, root names, nesting checks. It is *not* needed for the no-overwrite rule, where the filesystem's own case-insensitive `stat()` already @@ -2952,7 +3036,7 @@ The bulk of the codebase is portable because the portability rules in §10 were treated as correctness from the start. What the port needed is registered as **W1–W9** (§13.7) and is done; packaging is built and awaits a clean-machine run. -Two Windows-specific design points worth stating here: +Three Windows-specific design points worth stating here: - **The node runs in one of three modes, chosen at install and switchable afterwards** from the Node page: only while the application is open (it starts @@ -3090,7 +3174,7 @@ be understood, not so the incident can be retold. | **NS4** | **Operator authority comes from the node's roster and from nowhere else.** No auto-pin from the keystore, no resolution through the hub, no config key — a config naming one is warned about and never obeyed (§3.4, §6.1) | | **NS5** | The proof is **bound to the transport channel** (DTLS fingerprints / certificate hash), so a signaling relay that substitutes its own cannot produce it (§5.2) | | **NS6** | **`sender_id` is enforced from the authenticated session, never the wire.** It is what the store keys on; it is not what authenticates a message — the device signature is (§4.5) | -| **NS7** | The node authenticates to the hub with **Ed25519 and no password**, and its token's scope is refused for group management (§7.2) | +| **NS7** | The node authenticates to the hub with **Ed25519 and no password**, and its token's scope is refused for group management and on the admin and moderator API (§7.2) | | **NS8** | **The node refuses connections when it holds no group key.** There is no bypass switch | ### 13.3 Second review (code review) — the default numbering @@ -3118,7 +3202,7 @@ be understood, not so the incident can be retold. | **H4** | Revocation reaches nodes, drops live sessions, and **persists across a restart** (§7.5) | | **H5** | An admin challenge is a **structured, domain-separated transcript naming the operation and subject**, and the client refuses to sign anything that is not what the user asked for (§5.4) | | **H6** | Unauthenticated work a node will do is bounded: a small pre-handshake buffer, a transcode semaphore, per-user pending-offer caps, and a membership check on signaling (§7.2) | -| **H7** | **Only public groups register content hashes with the hub.** Private groups register nothing, and the swarm route requires authentication | +| **H7** | **No node registers a content hash with the hub**, for any group. The only content hashes it holds are of public content somebody reported (§7.5) | **Medium** @@ -3161,7 +3245,7 @@ be understood, not so the incident can be retold. | **M4** *(third review)* | Federation binds a pushed row's source to the signer, checks the token audience, caps the push, rejects replays, and scopes revocation to the peer's own entries (§7.6) | | **M5** *(third review)* | A CSP and security headers apply to the hub-served application, verified against the running app — a mis-tuned CSP shows as a blank page | | **M6** *(third review)* | **Withdrawn.** It misread the node registering a hub membership during the CLI invite flow — which is deliberate — as authorization drift | -| **L1–L11** *(third review)* | Opportunistic hardening: relay-registry proof of possession; delete orphaned modules rather than leaving them to be rewired; decide and document account enumeration; an aggregate upload quota; header-only control-API tokens; a freshness bound on revocation replay; state that the exact-hash content check is structural; validate group-name length and charset; require `exp` and bind an audience on token decode; key the rate limiter through the same client-address helper as everything else; keep diagnostic logging truncated | +| **L1–L11** *(third review)* | Opportunistic hardening: relay-registry proof of possession; delete orphaned modules rather than leaving them to be rewired; decide and document account enumeration; an aggregate upload quota; header-only control-API tokens; a freshness bound on revocation replay; validate group-name length and charset; require `exp` and bind an audience on token decode; key the rate limiter through the same client-address helper as everything else; keep diagnostic logging truncated | Two structural recommendations from that review stand as rules: @@ -3206,15 +3290,12 @@ the wrong thing** — a per-IP rate limit bounds a caller, never the mailbox tha receives what they cause, which is why `AV10` is a cooldown per *account* under a limit per IP rather than a tighter limit. -**Open for decision, not a defect:** `moderation.AUTO_BLOCK_THRESHOLD` is 3. -Three distinct accounts blocking a hash adds it to the list every node -enforces, network-wide, automatically, with manual admin removal the only -undo. That is already far better than the anonymous version it replaced, and -it is still a censorship primitive an attacker buys for the price of three -email addresses. Raising it buys little; requiring the reporting accounts to -be more than a day old would cost a patient attacker a day and cost an honest -reporter nothing after their first. Left as it is because it is a moderation -policy rather than a bug, and the person who sets that policy is the operator. +**Reports lead to a review, not a block** (§7.5). Distinct accounts reaching +the threshold used to add a hash to the list every node enforces, automatically, +with manual removal the only undo — a censorship primitive bought for the price of +three email addresses. Now the reporter must be a day-old account and a member of +the public group, and the threshold queues the hash for an administrator; blocking +without review is an instance setting, off by default. `admin.py` was read under this lens and needed nothing. Its moderator/admin line is drawn explicitly — a moderator may not change a role, may not revoke, @@ -3227,9 +3308,9 @@ had already been asked. | **AV1** | **An empty claim is a claim on nothing.** A node's group set is `authorized ∩ claimed`, and an absent or empty `group_ids` registers it for no group rather than all of its owner's — on registration and on `update_groups` alike (§7.2) | | **AV2** | **A client treats the hub's node list as candidates, not a ranking**, and tries the next one on a `not_hosted` refusal (`MESHBAY_NODE_PROTOCOL.md` §6.3) | | **AV3** | **A node speaks only for the groups it is registered for.** `chat_notify` names a group and is checked against that node's set before a notification is written for anyone, and it is rate-limited per node — the fan-out is one write per member | -| **AV4** | **Nobody names a third party's address.** A swarm source publishes a transport and a port, never a host; where a peer is comes from its node record, stamped with the address its announce arrived from. The number of hashes one account may claim is bounded | +| **AV4** | **Nobody names a third party's address.** Where a peer is comes from its node record, stamped with the address its announce arrived from | | **AV5** | **An answer is accepted only from the node the offer was sent to.** A `peer_id` is bound to its node, so no connected node can resolve another's pending offer | -| **AV6** | **A relay proves possession of its approved key.** A public key is not a password, and the register call is unauthenticated by design — it is not a user — so the proof is the only thing standing between a stranger and where nodes send relayed traffic | +| **AV6** | **A write that decides where other people's traffic goes proves possession of a key, never presents one.** A public key is not a password. The relay registry this was written for has been removed — nothing called it, and no TURN relay is needed (§11.1) — and the rule stands for whatever replaces it | | **AV7** | **A node bounds how many peers it holds and how long an unproven one lasts.** The hub's cap is per calling account, which is a limit on each member and not on the machine, so without this an operator's exposure grew with the size of their groups | | **AV8** | **One account cannot make the hub mail another at will.** The invitation email's subject comes from the group row, never from the request, and the endpoint is metered | | **AV9** | **No mail is sent from the event loop.** `smtplib` is synchronous and waits up to ten seconds; called from an async handler that wait is the whole instance's, not one request's. Every send goes through `mail.send_off_loop`. **Argon2 is held to the same rule**: every derivation runs on one dedicated worker thread (`auth.*_off_loop`), never on the loop and never two at a time, because two concurrent `lanes=4` derivations deadlock in OpenSSL. **So is the node's disk**: every filesystem call on a group's content — the stat as much as the read, since a stat is what wakes a sleeping disk — goes through `roots.off_disk`, onto one worker thread per root set. A spun-down or network-mounted root answers its first syscall in seconds, and on the loop that is every group, every stream and the hub socket waiting for a platter. **ffmpeg's own output too**, through `asyncio.to_thread` rather than that per-root thread: a temp file is not a group root and has no platter to serialise against, but a whole transcode read inline is still tens of megabytes of blocking read | @@ -3242,7 +3323,7 @@ had already been asked. | **AV19** | **Nothing carries the path to the migrations.** `meshbay-hub migrate` derives it from the installed package, so the RPM, the DEB, a venv and a checkout all agree. A unit naming `alembic.ini` names a file whose `%(here)s` stops being true the moment packaging moves it | | **AV18** | **The hub runs on exactly one worker, and says so at startup.** `_connected_nodes`, `_node_groups`, `_webrtc_answers` and the relay registry are per-process: a second worker makes a node intermittently unreachable for half its members, which is a symptom that describes something else entirely | | **AV14** | **MHP binds its audience, and the hub reads its own identity at call time.** A token is minted for one peer and accepted by that peer only. `federation.py` bound `_hub_id` and `_hub_sk_pem` at import, which is before `load_hub_keypair` runs, so it signed with `None` and called itself `meshbay.org` whatever the instance was named — and the verifier named no audience for the `aud` the issuer sets, which PyJWT refuses outright. MHP could not complete one authenticated request between two hubs | -| **AV15** | **A hash is checked for shape before it is a key lookup**, on the unauthenticated blocklist endpoints a node consults | +| **AV15** | **A hash is checked for shape before it is a key lookup**, on every blocklist endpoint, the administrator's included | | **AV20** | **Chat is bounded in size and in rate, like every other member-supplied write** (§6.6). A message is a row on the operator's disk that nothing expires, a relayed copy for every connected member and a notification for every member of the group; the only ceiling was the frame size. Uploads had carried four protections and a cap since C5a because somebody asked what one member costs the others on that path, and nobody had asked it on this one | | **AV21** | **A lease is what the node granted, not what the client called it** (§5.5). `tr` was read as a boolean, so any non-empty string skipped the leaseless ceiling and every cap behind it, and a queued transfer was held back only by the honesty of the client waiting in the queue | | **AV22** | **The node's own controls take no authority from a hub token** (§6.7). `node_status`, `node_settings_set`, `roster_read`, `denylist_read`, `denylist_clear` and `node_reload` were gated on the account id in the JWT, which is the hub's to choose — NS4 and M3 with the check written the other way round. The gate is a proved operator device, which a hub holding no user keys cannot produce | @@ -3254,6 +3335,7 @@ had already been asked. | **AV29** | **An invitation link is bounded on both halves and its mail on the sender** (§3.4, §7.3). Twenty outstanding per group on the node (bearer codes) and on the hub (tickets); and because a link mail reaches an address the hub has no relationship with, at the request of anyone who owns a group, it is counted **per sending account per day** (`mail.invite_link_daily_cap`, 10), under the recipient and instance bounds and outside the recovery reserve (`invite_link` is not a recovery purpose) | | **AV28** | **How many node keys one account may announce is bounded** (§7.2). Each is a row plus an IP-log row under a one-year retention, so an account in a loop writes a year of storage on the operator's disk having paid only for signatures. Proof of possession (**M8**) settles whose key it is and not how many. Counted only where a row is added: re-announcing a key already held keeps working at the ceiling, or a node that reached it could never refresh its address again | | **AV30** | **What one member's offers cost a node is bounded per account and per node, and the bound admits the heaviest ordinary account** (§7.2). Each offer makes the node allocate a peer connection. A budget of 120 per node refilled at two a second bounds a member there without touching their other nodes, and it is counted by account because a mobile carrier shares one IPv4 address among many subscribers. Pending offers are capped at 32 per account. Both refusals carry `Retry-After` and the client retries them, because a refused offer otherwise reads as a node that is down | +| **AV31** | **What waits for a signature is bounded** (§5.4). Any authenticated member can ask for an admin challenge, since the signature is checked afterwards, and a pending challenge kept its whole request until answered — measured, 200 requests of 1 MiB held 400 MiB for the life of one connection. At most eight pending per connection, 64 KiB each, expired ones dropped | ### 13.6 Chat design findings @@ -3274,7 +3356,7 @@ had already been asked. |---|---| | **W1** | Platform directories: no hardcoded XDG paths | | **W2** | Signal handling is platform-guarded | -| **W3** | Daemon lifecycle: a per-user startup launcher by default, a scheduled-task service mode offered, switchable after install (§11.2) | +| **W3** | Daemon lifecycle: three modes — only while the application is open, at sign-in, or as a boot-time scheduled task — chosen at install and switchable after; one start/stop implementation, the CLI's (§11.2) | | **W4** | Packaging: one per-user installer carrying client and node, with media tools bundled | | **W5** | File permission calls are skipped where they have no meaning | | **W6** | Media-tool discovery fails at startup with a stated reason rather than at first use | @@ -3422,9 +3504,7 @@ process runs it — `systemctl --user` on Linux, Task Scheduler on Windows. | **A signed upload transcript** | Ownership is recorded by the node and verifiable by nobody else (§5.4). Making it provable is a transcript the uploader signs, stored with the entry — designed in outline, not built | | Forward secrecy in group chat | **Given up deliberately and on the record** (§4.5). If it becomes a requirement it belongs in 1:1 DM | | Metadata at the hub | Membership, and who posted in which group and when. A known leak, not a solved problem (§7.1) | -| The exact-hash content check | Structural, not functional (§7.5) | -| **QUIC** | Off by default, and **not at parity**: it serves the index and file chunks with no transfer lease, no leaseless ceiling and no root-availability check, does its file I/O on the event loop, and returns exception text to the peer (**L3**). No client speaks it. Either it comes to parity or it goes; until then §5.1's "chat is the only gap" is the one sentence here that overstates the code | -| **The relay registry** | **Closed in the code**: `relay.RELAYS_ENABLED` is False and every `/v1/relays` route answers 503, as federation does. Nothing in the tree calls them, node or client, and §11.1 measured two ISPs with no TURN relay needed. Kept code that nothing calls is what **L7** says not to keep; it stays only as the proof-of-possession design (**AV6**) until a node needs a relay or it is deleted | +| **QUIC** | Off by default, and **not at parity**: it serves the index and file chunks with no transfer lease, no leaseless ceiling, no root-availability check and no content blocklist, does its file I/O on the event loop, and returns exception text to the peer (**L3**). No client speaks it. Either it comes to parity or it goes; until then §5.1's "chat is the only gap" is the one sentence here that overstates the code | | **A very high bitrate wedges the player against a small buffer ceiling** | Where even the *floor* read-ahead does not fit — ninety seconds plus the minute kept behind, at the file's bitrate, above what the engine will hold — the film stalls: measured on the harness at 9.3 Mbit/s against a 100 MB ceiling, 100.8 s of film played in 900 s of wall clock. **Predates the byte budget and is unchanged by it**, to the tenth of a second; what the budget did change there is the refusal count, 1560 → 2. The fix is not a bound at all, it is a second stage of buffer outside the SourceBuffer, which means gating the append path — the riskiest change in this area and not one to make alongside another | | **The reconnect backoff only wakes on `visibilitychange`** | So a tab that stays visible through an outage — which is what a screen wake lock guarantees while a film is playing — waits out the full backoff, up to 30 s, after the network is already back. Nothing listens for `online` | | **Per-device revocation has no CLI** | A device is revoked over MNP (`roster.revoke_device`), from a device the node has already pinned. On a headless node the operator's only lever is `member unpin`, which removes **every** device of that account — so the per-device control the roster is built around is reachable from an interface and from nowhere else. §6.7 listed a `meshbay-node member device list\|revoke` verb that was never written, and that listing is how this was found: `USERGUIDE.md` was the first document written by reading the CLI rather than this specification, and the verb it copied out did not run | |