aboutsummaryrefslogtreecommitdiffstats
Commit message (Collapse)AuthorAgeFilesLines
* feat(node): meshbay-node group add — host another of your groupsChristophe Besson2026-08-155-12/+224
| | | | | | | | | | | | | | | | | | | | | | | | | | | | Attaching a group to a node meant hand-editing node.toml with a UUID copied from a browser URL, restarting, and knowing that gek-init exists. Nothing in the CLI said so, and on a node reached over SSH there is no paste buffer to carry a UUID across in the first place. meshbay-node group add grenet --dir ~/grenet-share The name is resolved against the operator's groups on the hub by the daemon, which is the process holding the session. The [[groups]] block is appended to node.toml as text rather than round-tripped through a TOML writer: the file is hand-written and its comments explain decisions worth keeping. The directory is created, and the command says what remains — restart, then gek-init for that group. It refuses a name it cannot find by printing the groups it can, with their ids. That listing is the useful half of the answer and it was missing everywhere: _daemon_api now renders an `available` list from any endpoint that offers one. The key is per group and pairing is not, which is the part that reads as a gap until it is written down: one paired browser covers every group the node hosts, while each group's key admits only its own members. §4 of the user guide now says all three of those in one place. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(cli): --group takes a name, and operator pair says it takes noneChristophe Besson2026-08-151-2/+35
| | | | | | | | | | | | | | | | | | Two ways the CLI misled someone attaching a second group to a node. `meshbay-node operator pair --group grenet` accepted the flag and ignored it: pairing is node-wide and always was. That invites exactly the wrong reading — that a code belongs to a group, and that pairing had failed because the group did not change. It now refuses the flag and says one paired browser covers every group the node hosts. `--group` also only ever accepted a UUID. A name went through untouched and the daemon answered as though the group did not exist, which is not what happened. It now resolves a name against node.toml, and when there is no match it prints the groups there are, with their ids — the missing half of the answer. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* perf(upload): several chunks in flight, instead of one per round tripChristophe Besson2026-08-153-32/+120
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The uploader read a 48 KB slice, sent it, and waited for the node to acknowledge it before reading the next one. That caps throughput at one chunk per round trip regardless of available bandwidth, and it is worse than the arithmetic suggests: the sender is idle for almost the whole time, so SCTP's congestion window never opens either, and the transport stays slow even when the link is not. Measured against the real node over a 100 ms path (netem on loopback): 48 KB chunks, one at a time 0.16 MB/s 48 KB chunks, 32 in flight 3.47 MB/s On loopback with no latency both are ~32 MB/s, which is why nothing here ever caught it: the local end-to-end run cannot see a round-trip problem. transport.uploadFile() now keeps a window of chunks in flight and matches acks by arrival, with the node's own ordering rule as the guard — a DataChannel is ordered and reliable, and the node refuses any chunk that is not the one it expects next. It pauses when the channel's buffered amount gets high, so the progress bar keeps reporting what the node has taken rather than what the browser has queued. Both callers, the Files panel and chat attachments, go through it. The end-to-end harness grew an opt-in benchmark behind MESHBAY_BENCH=1 that removes its own files afterwards, and it taught me something about the harness rather than the code: it took an unsolicited index_sync push for an upload ack, because unlike app.js it had no place to put one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(groups): editable description, and one source of operator authorityChristophe Besson2026-08-1513-66/+348
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A description could only be set the moment a group was created, so every group made before anyone thought of one stayed blank for good. The owner can now edit it from the group's page, and PATCH /v1/groups/{id} takes it. That endpoint takes the description and nothing else, deliberately. The name, the visibility and the join policy are the terms members joined on; a private group that can quietly become public is not the group they agreed to be in. Changing those needs a decision about who gets told, not a field on a form — there is a test saying so. Separately, the legacy operator key is gone. `admin_pk_ed25519` in node.toml named the operator before the roster existed and was kept so that an existing deployment would keep working; nothing uses it, and a second source of node authority is not something to carry around out of politeness. Authority is the roster, read fresh on every check. It is removed rather than ignored: a config that still names the key gets a warning at startup pointing at the file. Dropping it in silence would refuse invites and file deletion with a signature error that looks like a bug somewhere else — which is exactly how finding M3 presented. Two tests were verifying admin operations by naming a key in the context, which was the legacy path. They now pair an operator into a roster, the way an operator does. The authority test anchored on the deleted function and passed vacuously once it disappeared; it states the invariant against the verifier and the daemon instead. Also defined .btn-secondary, used in four places and styled in none. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* The Files panel stops guessing, and search says how old its answer isChristophe Besson2026-08-153-32/+78
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | Four small things, two of which were the same bug wearing different hats. The index cache seeded the Files panel and then raced the live index: IndexedDB is async, so a fast node could hand you the real list and have it overwritten a moment later by the cached one. That is the "choses bizarres". The panel now shows what the node says, or says it cannot reach the node — no third state that looks like data but is memory. The cache stays, written on every sync, and the search page is the only thing that reads it. Search across groups has no other source: it cannot connect to every node to answer a keystroke. So it now says how stale each hit is — "synced 2 hours ago", per group — and a line under the results explains that opening a group refreshes what search knows about it. Deleting a file also rewrites the cache now; it used to refresh the table and leave the cache holding a file that no longer existed, which is why search kept offering it. The invite form is shown only once this browser holds an operator key. Invites are signed with it and the node checks the signature against its roster, so an unpaired browser could fill the form in and fail on submit. An owner who is not the node's operator is told to ask the one who is. Group descriptions now show on the home cards, as they already did in Explore. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(logs): keep the username on records the account no longer answers forChristophe Besson2026-08-158-4/+90
| | | | | | | | | | | | | | | | | | | | | The connection log took the name from a join on `users`, and deletion tombstones that row — so every record belonging to a deleted account reported `deleted-3f9a1c`, which is the one answer that helps nobody. The log is kept for a legal retention period precisely so it can say who did what; losing the name at deletion kept the data and lost the point of it. `ip_logs.username` is written as the account is erased, and stays NULL while the account is alive, where the join is better because it cannot go stale. The admin view prefers the stored name when there is one: the join still answers after deletion, just with the tombstone. Releasing the username for re-registration and keeping it in the log are separate things, and the guide now says so. On the node side, the pre-proof audit line records the username the session already knew, instead of leaving the column empty. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* fix(members): restore the member list and invite form, and retire the ↵Christophe Besson2026-08-154-48/+106
| | | | | | | | | | | | | | | | | | | | | | | | | | | pairing form Moving the invite form above the member list cut both out of MembersPanel and pasted them into AdminPage, where `doInvite`, `members`, `adminId` and `inviteCode` do not exist. A standard member saw an empty Members tab, the group owner saw only a pairing form, and the hub's own Users tab referenced four undefined names. The pairing form outstaying its welcome is a second bug and an older one. `is_node_admin` compares the connecting account with the account that owns the node — it says nothing about whether *this browser's key* was ever paired, which is the thing pairing changes and the thing that lets you sign an invite. So the form showed for an operator who paired months ago, accepted a fresh code, reported success, and stayed exactly where it was. The node already reports the roster role in `join_result`; the transport keeps it, and the form appears only when this identity is not an operator key yet. Also dropped a clause from the pairing hint: the code never passing through the hub is worth saying, the theory behind it is not. test_spa_ordering.py gets three checks for this class of bug — a cut-and-paste between components is invisible to every other test we have. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* docs: account deletion, notifications, and the APIs that no longer existChristophe Besson2026-08-154-124/+232
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Account deletion is the headline, in the user guide and in draft-v5 §6.1, and the important half is what deletion does *not* do. It releases the username, clears the email and password hash, drops memberships, notifications, refresh tokens and node registrations, and refuses any access token still inside its hour. It does not touch a node: files, the pinned identity and the keypair bundle stay on machines the hub does not command, which is the same sovereignty §5.5 relies on — so deleting a hub account is not an erasure request to the operators hosting you. The IP log survives too, attributable, for its legal retention period. The claims table in §2 gets a row saying exactly this, adversary by adversary. Notifications get a section: one entry per conversation rather than per message, never one for your own message, invitations that clear when you join, muting that lives on the hub so it works from any browser. Then the corrections, which is most of the diff. The guide still described a node HTTP API — `GET /index`, `GET /file/{id}`, an HLS playlist, and a `player.js` that does not exist — with curl examples inviting the reader to expose port 19001. That surface was removed in 0.2.0 as findings C1 and C6, precisely because it served files outside the handshake that decides what a peer may see. Sections 6, 7 and the API reference now describe MNP message pairs, and the quickstart says the same in French. Also corrected: the JWT table advertised a `pk_user` claim that no longer exists (it was what let the token issuer decide who could delete a file), `/pubkeys` no longer returns identity keys, and the GEK-distribution endpoints are gone entirely rather than merely unused. draft-v5 §5.2 had uploads landing in `.uploads/{user_id}/`; they land in `uploads/`, chat attachments included. §6.1 now says the hub learns the author's user_id from chat_notify — a stable identifier, and a metadata leak worth naming rather than leaving as "by whom". CLAUDE.md records why the deployed hub broke this week: create_all() creates missing tables, never missing columns, so a schema change passes every test (fresh DB per run) and never reaches production. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* One stroked icon set, matching the Administration shieldChristophe Besson2026-08-152-23/+94
| | | | | | | | | | | | | | | | | | | | | | | | | | | | The site drew its icons with colour emoji, which means a different drawing on every operating system and never the same line weight twice — except the Administration entry, whose shield is a text-default glyph and so rendered as a thin outline. That one was the odd one out and also the one that looked right. Fifteen inline SVG paths now replace it: menu, bell, shield, globe, gear, sun, moon, power, lock, envelope, door, clip, check, chevron, close. They are stroked in currentColor and sized in em, so an icon takes the colour and size of the text beside it and needs no rule of its own — the caret in the dark nav bar and the gear in a light dropdown are the same file. The explorer keeps its emoji. There the icon says what kind of file this is and the colour is doing real work; a wall of identical grey outlines would be a worse file list. One bug fell out of the change: .nav-notif was painted with --text, the page's text colour, on a nav bar that is nearly black. The emoji bell carried its own colours and showed anyway; a stroked one was invisible until it was given --nav-text. Checked by rendering the real stylesheet in headless Chrome, light and dark. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* Notifications: one per conversation, none for your own messagesChristophe Besson2026-08-1410-27/+422
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Four things were wrong, and they compounded: a busy chat produced one row per message, muting a group did nothing at all, there was no way to clear the list, and the one person guaranteed to know about a message — its author — was told about it. The author bug was a name mismatch across two processes. The node sent chat_notify without saying who wrote the message, so the hub used the node's own token subject, which is the operator's account. The skip therefore matched the operator and no one else: everybody was notified of their own messages, and the operator was notified of nobody's. The node now names the author and the hub reads that field. Muting lived in the browser's localStorage and nothing ever read it, so the checkbox was decoration. It is a column on group_members now, checked where the notification is created — a notification nobody wants is not written at all. Chat keeps a single row per (user, kind, group) whose date moves and whose read flag clears, so a conversation is one line saying when it last spoke. Clicking it opens the group and dismisses it; joining a group dismisses its invitation; and DELETE /v1/notifications clears the lot. The hub deploy now runs alembic. create_all() only creates missing tables, so group_members.muted never arrived on the running hub and /v1/groups/mine answered 500 — worth catching in the script rather than in a browser. Verified end to end against the deployed hub and node: the author receives nothing, the other member receives exactly one, carrying its group_id. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(account): a user can delete their own account, an admin can delete oneChristophe Besson2026-08-147-5/+402
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Both go through the same erasure, so there is one description of what happens rather than two that drift. Gone: credentials, email, node key, group memberships, notifications, refresh tokens, node registrations. The username is released. Kept, on purpose and stated in the UI: the row itself, emptied, and the IP log that points at it. Those logs exist for a year to answer legal requests, and a log that can no longer say whose connection it recorded keeps the data while losing the only thing it is for. So the account becomes a tombstone rather than a hole in the table. Out of reach, also stated: files uploaded to nodes, and the identity keys nodes pinned. Those are on machines the hub does not command, and only their operators can remove them — `member unpin` and a delete on their own disk. Saying so in the confirmation matters more than the button. Owning groups blocks deletion, with the list. Cascading would delete other people's groups out from under them; the account holder can hand them over or delete them first, deliberately. Self-deletion re-checks the passphrase. A live token may be a borrowed laptop or a tab left open, and it is not consent to something irreversible. Admin deletion requires admin rather than moderator: suspension is the reversible moderation tool and stays one click away. A deleted account's access token stops working at once — the status check already refuses anything but "active", which the tests now pin down, because refresh tokens being gone would otherwise leave up to an hour of usable session. Tests: 8 covering what survives and what does not, plus a db_session fixture for assertions that cannot honestly be made through the API. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(files): one uploads/ directory, for files and chat alikeChristophe Besson2026-08-144-66/+107
| | | | | | | | | | | | | | | | | | | | | | | | | | Correction to the previous commit. Uploads went wherever the member happened to be looking, which spreads chat attachments through the tree and makes the destination a client-supplied path — surface that had to be defended. Everything a member sends now lands in `uploads/` at the root of the shared directory: visible, one place, easy for the operator to look into or empty. Chat attachments go there too, so the separate out-of-tree thumbs directory is not needed and is not built. They were already ordinary uploads; now they are ordinary uploads that land somewhere sensible. The destination is chosen by the node, so a client naming somewhere else changes nothing — the traversal surface simply is not there on this path. safe_subdir() remains for dir_create, where the path genuinely does come from the client, and keeps its tests. One shared directory means name collisions are ordinary rather than adversarial: every camera produces IMG_1234.jpg. The node finds a free name — "IMG_1234 (2).jpg" — and reports it in the ack, because a chat message has to point at the file that was actually written and not at someone else's. Nothing is ever replaced, which is the property the per-user quarantine existed for (C5a) and the one the tests assert; they fail if the free-name search is removed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(files): upload into the current directory, and create foldersChristophe Besson2026-08-146-76/+280
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The per-user quarantine is gone. `.uploads/{user_id}/` was the fix for C5a, and it worked, but it made the shared directory something nobody could organise: every file landed under a uuid nobody recognises. Files now go where the member is looking, most often the root. What the quarantine actually bought is kept, and is now what the tests assert rather than the location: - an existing file is never replaced. That was the real defect — overwriting a file also made the attacker its recorded uploader, and therefore able to delete it through the uploader path - the name allowlist is unchanged - the destination is confined under the shared root That last one is new surface: the directory arrives from the client. safe_subdir() is the single place that decides, with two independent guards — every segment against the name allowlist, and the resolved result under the root — because one of them will eventually be refactored by someone who does not know why it is there. Ten traversal cases are covered, and they fail if both guards go. Also adds `dir_create` (any member may organise a shared directory; audited like anything that writes to the operator's disk) and makes the node report its real directory list in index_sync — folders were inferred from file paths, so a new empty one, or one that had been emptied, simply did not exist as far as the UI was concerned. Two C5a tests changed their assertions deliberately, as C5b's did before: they encoded the quarantine path, which is the thing being removed. The property they existed for is asserted more directly than before. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* feat(ui): upload progress, PDF preview, invite form above the member listChristophe Besson2026-08-143-22/+69
| | | | | | | | | | | | | | | | | | | Three of the ten UI items, the ones that needed no decision. Upload showed a disabled button and nothing else — on a large file that is indistinguishable from a hang. It now has the same bar downloads have, filled from chunks the node has acknowledged rather than bytes read locally, and says "indexing" for the pause at the end: the node re-indexes on a filesystem event, so there is a real wait there with nothing to poll. PDFs open in the overlay through the browser's own viewer. The file is decrypted in the page as any other preview is, and shown from a blob: URL — nothing leaves the tab, and no external viewer is involved. The invite form sits above the member list, where the action is rather than after the thing it acts on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* chore: release 0.2.00.2Christophe Besson2026-08-1413-19/+35
| | | | | | | | | | | | All three packages together, as the conventions require, plus the RPM and DEB metadata and their changelogs. The tag said 0.2 while every package announced 0.1.0, which would have shipped an RPM claiming to be the reviewed build while containing a different protocol: the hub schema lost the user identity keys, tokens lost pk_user, and gek_bundle_store left the wire. Pre-1.0, a breaking change bumps MINOR. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
* merge: Phase 11.5 security remediation, invite redesign, per-node identityChristophe Besson2026-08-1463-2279/+8470
|\ | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Brings in the security remediation branch. Three bodies of work, and what they changed about what this project may claim. Phase 11.5 closed the gap between the documents and the code: the unauthenticated node HTTP API and the TCP transport deleted, one handshake shared by the remaining two transports, mutual authentication, structured admin transcripts, upload confinement, group isolation, revocation that reaches nodes. Six critical and seven high findings closed, bounded, or deferred by decision. The invite redesign closed H3 and M3 — the last open High. The hub was the key directory: an inviter fetched the invitee's key from it and wrapped the group key for whatever came back, so a hub answering with its own key was handed the group key by an honest member following the protocol exactly. That lookup is gone. The node holds the group key and wraps it itself, for a key its recipient proves possession of, bound to an account by a one-time code the hub never sees. M3 fell out of the same work: node authority comes from a local roster, never from the hub. Per-node identity cut what remains of C4 down to one operator. A single keypair used to be copied to every node its owner joined; each node now gets its own, so cracking the bundle on one machine yields a key that is a stranger everywhere else — and on that machine, one that unlocks nothing its holder did not already serve. The bundle KDF moved to Argon2id 128 MB, and the hub stopped storing or publishing user keys at all. What this project may now say: the hub cannot read your content unless it ships you malicious client code. T3 remains, accepted (D1), and is what the native client removes. C4 is reduced, not closed, until 13.3. Chat is still plaintext at rest until Phase 15. Draft-v5 §2 states each claim against the adversary it holds against, which is the convention this branch exists to keep. Four defects were found by deploying it and using a browser, none by the test suite: a node going deaf on its hub socket, a token that predated group membership, a client reading values before they were assigned, and identity keys a browser held but never re-read. The lessons are recorded in CLAUDE.md. Tests: 343 across the three packages, plus QE/deploy/e2e.py — register, pair, invite, join, download, stream, second browser, revoke — run against the live deployment on a wiped hub and node.
| * fix(node): an operator keeps their role when reconnectingChristophe Besson2026-08-141-1/+5
| | | | | | | | | | | | | | | | | | | | | | An operator's roster row is node-wide, so looking it up by the group they happen to be opening found nothing and the client was told it had no role on a node it administers. Falls back to the node-wide row. Surfaced by running the live workflow twice: the first pass pins, the second is recognised — and only the second exercised this path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * docs: user guide and conventions catch up with per-node identityChristophe Besson2026-08-142-18/+32
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | USERGUIDE said registration submits your public keys "so other members can wrap GEK bundles for you". Both halves are wrong now: registration creates an account and nothing else, and nobody wraps anything for a key fetched from the hub. The API reference and the register body followed the same correction. CLAUDE.md gains the block a future session needs before touching registration or anything shaped like a user's public key: keys are born at first contact with a node and stay there, the hub publishes none, tokens carry no pk_user, and a scripted signup is now a real account. Left alone deliberately: first-review.md, docs/poc-v1*.md and poc/spike-results.md still describe the old JWT and registration. They are records of what was true on their date, like second-review's verdict table, and draft-v5 is what states the present. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * feat!: identity keys per node — C4's blast radius drops to one operatorChristophe Besson2026-08-1417-303/+493
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | One keypair was copied to every node its owner joined, so cracking the bundle on any single node yielded the identity used on all of them: their content on other operators' machines, and the ability to sign as them anywhere. That lateral reach was the part of C4 worth attacking. Each node now gets its own keypair, generated the first time its owner joins it and left with that node alone. An operator who cracks what sits on their own disk holds a key that is a stranger to every other node — and on their own node, one that unlocks nothing they did not already hold: they serve the content, the index and every byte of it by design. Nothing changes for the user. A first contact with a node already needed that operator's code, and the key is created in the same step; a second browser still recovers it from the node with the passphrase alone. Two operators can also no longer tell they host the same person by comparing keys. BREAKING, and deliberately without a compatibility path — the deployment is wiped for the next demo: - users.pk_ed25519 / pk_x25519 dropped (migration a7c31f9e40b2) - registration no longer sends or stores a key - PUT /v1/users/me/keys and regenerateKeys() gone; rotation is now `member unpin` plus a fresh code, decided on the machine that pinned it - /pubkeys returns an account id and a node's linking key. It was the directory H3 read, and nothing wraps for it any more - the pk_user JWT claim is gone That last one closed a live defect the inventory turned up: the node recorded pk_user as the uploader's identity and authorized deletion against it, so a hub issuing a token naming its own key could delete anyone's uploads on any node. Attribution now uses the key the node itself pinned. A simplification falls out. Registration generates nothing, so a scripted signup is a real account: `demo.py bootstrap` takes a wiped hub and node to a working demo with no browser, which was impossible while keys were born in one. Also fixes, found by running it on a wiped deployment: the key handed back on a join now belongs to the group the connection is for, not the group named in the invitation — an operator pairs node-wide but redeems the code while opening a group, and expects to read it. Tests: 343, including the two that state the property — a key pinned by one node is refused at another, and someone else's code does not admit it. Verified end to end against a wiped hub and node: bootstrap, pair, invite, join, download, stream, second browser, revoke. Design: docs/per-node-identity-v1.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * docs: Argon2id, the multi-browser property, and what a browser foundChristophe Besson2026-08-145-37/+139
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | draft-v5 §7 rewritten around the keypair bundle, because that is where the last open finding actually lives. New §7.1 states the adversary (an operator holding their own node's disk), what cracking a bundle yields (identity keys, hence content on *other* nodes and the ability to sign as that user — not the content they host in the clear by design), and the measured numbers rather than adjectives: PBKDF2 241 ms vs Argon2id 88 ms natively, a GPU ceiling moving from ~8k to ~2k guesses/s, six days for a 10⁹ dictionary run, four random words outlasting the sun. The honest summary is in there too — a factor of four on one card, not a thousand; what it buys is the cost of scale. §2 gains the row the table never had: **your identity keys stay yours**, ⚠️ against a malicious node operator. An operator hosts your content by design, and that was documented; that they can also try to become *you* was not. That is the difference between reading what they host and reading what other operators host. §4 records that the challenge now carries `node_pk`, why (a first-time member signs a transcript naming the node and has no GEK to complete a handshake with), and that it is checked against the ack rather than trusted. Also that refusals carry a code, and what `not_a_member` usually means. §8.1 states the multi-browser property plainly — one identity across browsers, recovered with the passphrase, no second code — together with its cost, since it is the same mechanism as C4. invite-pairing-v1 is no longer "a proposal": it shipped. §9bis gains the four browser-found failures and their common thread — e2e.py is a second implementation of the client, written in the right order by construction, so it proves the protocol and nothing about app.js. CLAUDE.md gets the two things a future session must not rediscover the hard way: the KDF parameters live in three places held identical by a parity test, and an unbounded await on the hub socket makes a node silently unreachable (three found). second-review: C4 marked reduced, not closed. devel-phases-next: 12.2's CSP must keep `wasm-unsafe-eval`, or the strict policy locks every user out of their keys. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * perf(client): bundle KDF to 128 MB, and derive it once per sign-inChristophe Besson2026-08-145-14/+37
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Argon2id memory 64 → 128 MB. Memory is the lever, not time: it caps how many guesses a card can hold at once, so the ceiling on one high-end GPU moves from roughly 4k to roughly 2k guesses/s and its 24 GB fits ~187 lanes instead of ~375. Measured through the vendored build: 640 ms, against 322 ms at 64 MB. While measuring the real cost of a sign-in, found the SPA deriving the bundle key twice — once for the key pair kept for the session, then again inside decryptBundle() for the local bundle. At these parameters that is 0.6 s of pure waste. Measured now, end to end: auth_key (PBKDF2 600k) 239 ms bundle v1 (PBKDF2 600k) 240 ms legacy, until every bundle is upgraded bundle v2 (Argon2id 128MB) 650 ms ----------------------------------- sign-in 1 129 ms (889 ms once no v1 bundles remain) Once per sign-in, and only then: reopening a group, downloading, streaming and reloading the page all reuse the key, which lives in IndexedDB from login. Also bounds two waits in the node's hub WebSocket, found because the node went silent again mid-deploy. It had reconnected after the hub restart, sent its auth frame, and waited for a reply that never came — `ws.recv()` had no timeout, so a hub that accepts a socket and then says nothing for a few seconds while starting up parks the task forever: node running, logging nothing, invisible to everyone. The auth exchange now times out at 15 s, connect at 15 s, and a refused auth retries with a fresh token instead of ending the task for good. QE harness signs in once per account and reuses the token — several clients there stand for several browsers of one person, and what tells them apart is which keys they hold, not which token, while the hub quite rightly rate-limits repeated logins from one address. Tests: 341, plus the live workflow. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * feat(client): Argon2id for the keypair bundle, and remove the backup toggleChristophe Besson2026-08-149-100/+315
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Two corrections to yesterday's judgement, in the order they matter. **The toggle is gone.** Asked to make the remote key backup optional, I shipped a setting whose "off" position meant: no second browser, ever, and clearing your storage destroys the account. I wrote the warning that says so without drawing the conclusion. A control whose only effect is to break the ordinary case is not a control, and removing an exposure by removing the feature is not a fix. Every browser backs its keys up again, unconditionally. **The exposure is fixed where it actually lives: the KDF.** The keypair bundle rests on every node whose group its owner joins, protected by the passphrase alone (finding C4). It used PBKDF2-SHA512 at 600k — compute-only, which is exactly what a GPU eats. Measured on this machine: PBKDF2 600k costs 241 ms and Argon2id 64 MB/t=3 costs 322 ms, near enough the same honest work, except only one of them forces an attacker to find 64 MB per guess. So the bundle key is now Argon2id 64 MB / t=3 / p=1, via a vendored WebAssembly build (no external host — the CSP forbids one, and 12.2 will tighten it further). Parameters chosen by measurement through that build: 19 MB is OWASP's floor at 118 ms, 256 MB is 1.3 s and too slow for a phone, 64 MB sits where a login should. What this buys, stated honestly: cracking a bundle yields the owner's identity keys, and with them content on OTHER nodes and the ability to sign as them — not the content on the operator's own node, which they host in the clear by design. Argon2id raises that price steeply; it does not remove it, and a weak passphrase still loses. Hence the floor raised to 12 characters and ~60 bits in the same breath, which can only be enforced client-side: with the password split (T1) the hub never sees a passphrase. Migration is automatic and invisible. Bundles carry an "MBK2" marker; the old form is still readable, and is re-encrypted the first time a browser backs it up. Both keys are derived at sign-in, because which one a bundle needs is only known once it is read and the passphrase is deliberately not kept around. Two implementations of the KDF now exist — the browser's WASM and argon2-cffi in QE — so a parity test holds them byte-identical. A disagreement would not look like an error; it would look like an account nobody can open. keypair_bundle_delete stays, without a UI. It is the mechanism behind withdrawing your data from a node, exercised end to end, and it will belong to a deliberate "forget me on this node" action rather than a setting that quietly disables multi-device. Verified against the live deployment: the full workflow passes, including recovering keys on a second client from the passphrase alone. Tests: 341. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * feat(client): make the key backup a choice, and raise the passphrase floorChristophe Besson2026-08-146-11/+190
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Two things the multi-browser story made obvious. **The backup is now opt-out.** Keys are kept, encrypted with the passphrase, on every node whose group you join — that is what lets a second browser recover them, and it is finding C4: a PBKDF2-protected blob on other people's disks, attackable offline at the speed of PBKDF2, which is memory-light and therefore cheap on a GPU. Until now everybody paid that cost, including people who will only ever use one browser and get nothing back for it. Settings → "Use this account on other devices". Turning it off does not merely stop future uploads: the next connection to each node withdraws what that node already holds (new keypair_bundle_delete, which only ever deletes the caller's own, taken from the authenticated session and never from the message). The warning says plainly what it costs — clearing the browser then loses everything encrypted for that account, with no recovery, which is the point of choosing it. Default is on. Silent, unrecoverable key loss is worse for an ordinary user than an exposure the roadmap already tracks, but that is a judgement call and it is now visible and reversible instead of implicit. **Passphrase floor 8 → 12 characters, plus a strength estimate** shown while typing, with a refusal below ~60 bits. This number matters more here than in most applications: it is what stands between a node operator and your identity keys. It has to live in the client — with the password split (T1) the hub never sees a password and cannot enforce anything about one — so the UI says why it is asking, rather than nagging. The estimator is deliberately conservative and dependency-free: character classes and length, penalised for repetition and for the handful of patterns everyone tries. Verified against the live deployment: withdrawing the backup leaves a second browser unable to recover anything, which is exactly what it promises, and re-enabling restores it. Tests: 338. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * test: prove a second browser works after pairingChristophe Besson2026-08-141-0/+16
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The mechanism was already there — the encrypted keypair bundle goes to the node after a first successful connection, and any client holding the password can recover it — but nothing exercised it. e2e.py never pushed a bundle, so the case that matters to an ordinary user was the one case never tested. It now does what app.js does: backs the member's keys up to the node, then opens a second client carrying nothing but a username and a password. Against the live deployment that client recovers its identity keys, is recognised as the same person with no second code, gets the same group key, and browses the group. Also guards the ordering this depends on: the keypair bundle must be fetched before joinGroup() runs, or a browser that did not register has no key to sign the join with — invisible on the browser that did register, broken on every other one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * fix(client): recover identity keys from what the browser already hasChristophe Besson2026-08-141-0/+40
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Registering in Firefox and coming back to it said "This browser does not hold your keys" — while both halves of those keys were on disk a few bytes apart. _sessionKeys lives in sessionStorage, which dies with the tab. The encrypted keypair bundle is in localStorage from registration, and the key that opens it is in IndexedDB from login, but nothing ever put the two together again: only the login path did, and a returning user is restored from stored auth without logging in. So closing a tab looked identical to never having registered there. Recovery now happens before connecting: bundle from localStorage, key from IndexedDB, public half derived from our own secret rather than read back from the hub. The bundle is also queued for backup to the node, which is what lets a second browser recover the same keys with the password. Not hardening, and not from the invite redesign — an oversight in session restore that the redesign made visible, because joining is now the first thing that needs those keys. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * fix(client): capture the challenge values before joining, not afterChristophe Besson2026-08-142-8/+100
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | join_request signs a transcript over the node key and the node nonce, and runs before the GEK proof — a first-time member has no key to prove with. Both values were read further down, beside the proof that also uses them, so by the time joinGroup() ran neither was set and every invited member got "Handshake incomplete — reconnect and retry". They are now recorded the moment the challenge arrives. Third bug of the same shape found in a browser, and the reason is worth writing down: QE/deploy/e2e.py cannot catch any of them. It is a second implementation of the client, written in the right order by construction, so it passes while the SPA fails. It proves the protocol; it proves nothing about app.js. So this adds ordering guards over transport.js — source-level, which is not how one would normally test behaviour, but it is what sees this class of mistake: - node_pk and nonce_node are captured before joinGroup() runs - the join happens before the GEK proof - the ack still verifies the key the challenge announced Verified the way the suite requires: each fails against the source as it was, on the ordering assertion rather than on a missing marker. e2e.py also waits for the node to re-register rather than reporting "no nodes" at whoever just restarted the hub. Tests: 337 across the three packages. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * fix(client): refresh the token when the node says "not a member"Christophe Besson2026-08-145-13/+87
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A member added to a group after they signed in was refused by the node, told "Not a member of this group", and had no way forward but to log out and back in. The hub bakes `groups` into the access token at login and never pushes updates, so the token said they were in nothing while the database said otherwise. This lands on every newly invited member, at their first action, and the message tells them the opposite of the truth — toto2 was a member of newdemo on the hub and read that they were not. The refusal now carries a code the client can act on (`not_a_member`) rather than prose it would have to string-match, and the SPA refreshes the access token once and retries. Refreshing re-reads membership from the database, so the retry succeeds. Once per mount: if a fresh token still says not a member, that is the truth and it gets shown. The SPA had stored a refresh token since Phase 8 and never used it. It does now. Found in a browser, doing the ordinary thing — the automated run never sees it, because e2e.py logs in after being added to the group. Tests: 233 node+common, including a handshake test that the refusal carries the code, and the full e2e run against the live deployment. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * fix(node): keep reading the hub socket while negotiating WebRTCChristophe Besson2026-08-142-11/+67
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A node could be running, healthy in its own logs, and invisible to the hub with nothing to say why. That is what "No nodes available" looked like from a browser, and restarting the daemon was the only way out. maintain_ws awaited the WebRTC offer handler inline, inside the loop that reads the hub socket. One negotiation that did not finish — a client that closed its tab mid-ICE is enough — stopped the node reading that socket at all: pings unanswered, close frame never seen, later offers never served. The socket sat in CLOSE-WAIT with the hub's goodbye unread in the receive queue, which is how this was finally pinned down. Offers are now answered in their own task, so the read loop keeps draining whatever happens to any one peer. With that in place the existing reconnect logic works: a hub restart is seen (1012), retried through the 502 while it comes back up, and reconnected unattended — 19 seconds in the run that verified this. Also: - explicit ping_interval/ping_timeout. This connection is how a node stays reachable, and a half-open socket looks exactly like a working one. - a clean close ended `async for` without raising and reconnected in silence; it now says so, because a node that stops being reachable should leave a trace. - a failed negotiation logs the peer instead of taking the loop down with it. Predates this branch (Phase 11), and independent of the invite work — surfaced while testing it, because deploying the hub mid-session is exactly the trigger. Tests: 232 node+common, plus the full QE/deploy/e2e.py run against the live deployment after a deliberate hub restart. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * fix(node): announce the node key in the challenge, and keep names in the rosterChristophe Besson2026-08-147-32/+105
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Both found by deploying the thing and running the workflow end to end. Neither was reachable from the test suite, for the same reason in each case: the tests knew something a real client cannot. 1. A first-time joiner had no way to learn node_pk. join_request signs a transcript naming the node, and the node key was only sent in handshake_ack — which an invited member cannot reach, having no GEK to prove. joinGroup() therefore threw "handshake incomplete" and the browser path for an invited member was broken. Every test built the transcript from a node key it already had, so nothing noticed. The challenge now carries node_pk. It is unverified at that point and never a substitute for the ack: the ack still proves possession and signs the transcript, the client checks the two values match and refuses a peer that changed identity mid-handshake, and TOFU pinning is unchanged. A wrong value only makes our own verification fail. test_invite_then_join_delivers_the_gek now takes the key from the challenge instead of from sk_node, so it proves a real client can learn it. 2. The roster pinned everyone without a name. `_do_join_request` took the username from the session, which takes it from the JWT — and the hub puts no username claim in a token. So identities were pinned with an empty name and `member revoke <name>` could never match: the live node answered "known: , ,". Invitations now carry the name (new invites.username column, with a migration for the roster DBs already out there), and the CLI resolves a name through the daemon: its own roster first, the hub as fallback for identities pinned before this. The harness that found them is QE/deploy/e2e.py — gitignored with the rest of QE/, so it is not in this commit. It does the SPA's job in Python against the live deployment: hub login, WebRTC via hub signaling, the unified handshake, joining with a code, index, chunk download and MSE segments. Verified against meshbay.org and the local node: an account registered from scratch is invited by code, receives the group key wrapped for a key it proved it holds, downloads and decrypts a file, streams 5 encrypted fMP4 segments, reconnects with no code, and is refused after `member revoke`. The node audit log shows invite_create → join_pinned(via=code) → gek_wrapped → handshake, then join_no_gek once revoked. Tests: 232 node+common. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * docs: record the invite redesign — H3 and M3 closedChristophe Besson2026-08-145-98/+226
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | draft-v5 §2: against an active hub, reading content moves from "❌ H3" to "❌ T3 (browser) · ✅ native". The defensible sentence becomes "the hub cannot read your content unless it ships you malicious client code" — T3 is now the only path, it is an artifact rather than a silent directory lie, and it does not exist for a native client. New §5.5 describes admission and key delivery, with the four properties that carry it and the one exception (open-join groups, where the hub can walk in the front door — a property of open joining, and the setting is read from node.toml). Corrected while writing it: §5.1 said the C5b fix stopped a group admin who does not run the node from inviting, and that the redesign reverses this. It does not, because delegation was deferred. What changed is the timing — the operator issues a code and is then out of the loop. devel-phases-next: 12.1 is done and NOT as written. The plan was key transparency plus safety numbers; what shipped removes the directory read instead. Safety numbers make substitution detectable by a human who checks, at first contact, when there is nothing to check against. 12.2 (served-SPA integrity) is now the highest-value item in that phase. Phase 14 marked for what landed. second-review: H3 and M3 annotated closed at the finding, with what actually closed them. The §7 verdict table is left intact — it is the record of an audit on a date, and falsifying it would be worse than leaving it — with a note pointing at draft-v5 §2 for current state. CLAUDE.md matters most here, being loaded every session: NS4 read "admin_pk_ed25519 auto-pinned from keystore ✅ DONE", which is M3 described as a feature. Rewritten, with the two fixes that must never be attempted (auto-pin, hub lookup). QE/deploy/README.md: set-admin-pk retired from the walkthrough; the regression checklist now exercises pairing, joining by code, recognition without a code, and revocation. USERGUIDE.md is beyond the invite work but was actively wrong: it told users to POST GEK bundles to a hub endpoint deleted in Phase 12, and to re-wrap for every remaining member on revocation. Both replaced with what the code does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * feat(node): operator surface — member list, invite, revoke, unpin over SSHChristophe Besson2026-08-147-14/+499
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A node admits people from its own roster, and until now a headless operator had no way to put anyone on it: pairing worked from the CLI, everything else needed a browser on a machine that does not have one. Absorbs milestones 14.3/14.4. member list who is admitted, role, status, when and how pinned member invite <username> one-time code; the node wraps the key when they connect, so nobody has to be online then member revoke <username> stop serving them the key member unpin <username> forget the pin so they can pair again after a reset All of it goes through the daemon's loopback API with the per-run session token (11.5.3) — _daemon_api() in daemon.py, which also replaced three hand-rolled urllib blocks. `status` deliberately still reads the keystore, config and roster directly, so it works while the daemon is stopped. Two things the commands say out loud, because getting them wrong is silent: - revoke ends by telling the operator to rotate the key. The ex-member stops receiving it on their next connection, but they hold the current one, and "revoked" reads like it took the key back. - revoke/unpin refuse a username the roster does not know instead of acting on nobody. A typo must not look like success. Code lifetimes now differ by what the act is: 7 days for an invitation, which crosses a human conversation and gets answered whenever someone reads their messages, and 24 h for operator pairing, which is typed during the SSH session that printed it. Both configurable ([node] invite_ttl_hours, pair_ttl_hours). A day was long enough for the second and not for the first — a code that dies over a weekend means finding a browser to issue another one. The roster is also in the local admin UI, escaped: usernames come from the hub and land on the page that can re-key groups and read the audit log, so H2's rule covers them exactly as it covers filenames. Verified by driving the real CLI against a stub daemon over a socket, which is how the "known: <nothing>" bug in the not-found path turned up. Tests: 89 node here (roster, endpoints, CLI routing, TTL config). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * feat(node)!: the node wraps the group key — closes H3 and M3Christophe Besson2026-08-1418-285/+2637
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The invite flow fetched the invitee's pk_x25519 from the hub and wrapped the GEK for whatever came back (app.js:1466, and gek-init did the same server-side). The hub is the key directory, so a hub answering with its own key was handed the group key by an honest member following the protocol exactly. No forgery, no injection, nothing for the client to notice. That was H3. The fix is not safety numbers. Nobody reads the directory any more: - the node holds the GEK and wraps it itself, on every connection, for the X25519 key the joiner signed with their Ed25519 identity in one transcript (meshbay:join:v1), so the identity key vouches for the encryption key; - identities are bound to accounts by a one-time code the hub never sees — 40 bits, single use, one account, bounded per connection AND node-wide; - the node's own roster decides who may receive the key. Hub membership lets someone reach a node; it no longer gets them anything. A hub that invents an account and mints it a token is answered not_authorized_for_group. Safety numbers would have made substitution detectable by a human who checks, at the moment there is nothing to check against — first contact. Removing the lookup makes it impossible, and costs the user one code to pass along. M3 falls out of the same work. The daemon auto-pinned its own keystore key as admin_pk_ed25519 while the browser signs with the user identity key, so every privileged operation failed closed with a signature error that looked like a bug somewhere else; the demo only worked because a deploy script overwrote the value. Authority now comes from the roster, established locally by `operator pair`. Asking the hub for the operator's key — the obvious-looking fix — would have let the hub install itself as node administrator. BREAKING: gek_bundle_store is deleted, not gated. No member hands the node key material at all, so C5b becomes structural rather than an authorization to check. Existing stored bundles are still served, so current deployments keep working. Also: - join_policy (invite|open) is read from node.toml, never from the hub — a hub able to declare a group open would be handed its key. Unknown group ⇒ invite. - admin signatures are verified against the roster on every check, so unpinning takes effect without a restart. admin_pk_ed25519 stays readable as legacy. - two C5b tests were rewritten, deliberately: they asserted that gek_bundle_store demanded an operator signature, and the message is gone. They now assert the stronger property. The file says not to fix these tests, so this is the record of why they changed. - a slice-1 bug found while writing slice 2: connect() never passed skEdB64, so pairing would have failed at runtime with no test able to catch it. Tests: 152 node+common here, including an end-to-end DataChannel run where a member who has never held the group key redeems a code in the pre-proof window and receives the key wrapped for a key only they can open. Design: docs/invite-pairing-v1.md Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * test(node): allocate daemon test ports instead of hardcoding 28000Christophe Besson2026-08-141-3/+17
| | | | | | | | | | | | | | | | | | | | | | | | test_daemon.py started the real admin UI on a fixed port, so every test file that also brought up a node collided with it. Each file passed on its own and the full node suite failed with EADDRINUSE on test_daemon_creates_chat_store — which reads as a flaky regression rather than a test-isolation bug. Confirmed against a clean worktree at HEAD before touching anything: the failure predates the invite work. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * feat(node): CLI for headless operators — status, ui, gek-initChristophe Besson2026-08-133-10/+153
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Every operator action lived behind a web UI on the node's own loopback interface. For the normal deployment — a node on a server reached over SSH — that is unusable: no browser on the host, and 11.5.3 added a per-run token that had to be copied out of a log to get in. status hub, node public key, daemon state, groups, admin-key pinning. Reads the keystore directly so it works while the daemon is STOPPED, which is exactly when it is needed: the daemon cannot stay up before its key is linked or before a group exists. ui prints the URL and the ssh -L line. It does not open a browser — that was an assumption about the environment, and a wrong one. gek-init initialises a group key through the daemon's loopback API. Same operation as the admin UI button, no browser involved. Also fixes a latent bug in QE/deploy/deploy-node.sh: the pkill pattern was unanchored, so it matched any shell whose command line merely mentioned the daemon — including the one running the script. It killed a session three times before being pinned down. Anchored to the end of the command line. Verified against the live deployment. grenet and cbesson both connect over WebRTC through real NAT and can browse, download, stream, upload and chat. The node audit log confirms the security properties in production: uploads land in .uploads/{user_id}/ (C5a), the invite required the operator's signature over an admin transcript (C5b, H5), the pre-proof bundle window is bounded and audited (C4), and a non-member handshake was refused. Docs updated: Phase 14 marked partially delivered with the reason, draft-v5 §5.3 records the two operator personas, QE/deploy/README.md documents the commands and the remaining browser-only gaps (invite, delete). Tests: 121 node. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * feat(node): keep the daemon alive when the hub rejects credentialsChristophe Besson2026-08-131-8/+22
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | A node whose owner has not registered yet got a plain 401 from /v1/nodes/auth, which _login_with_retry re-raised — so the daemon exited and took its local admin UI down with it. That UI is where the operator reads the node's public key in order to link it, so exiting strands them: no daemon, no key, no way forward without digging the keystore open by hand. The daemon already parks on "No node key" for exactly this reason; it now parks on any 401, reporting waiting_for_account with a message naming the account and hub, and keeps retrying every 30s. The intended order remains: register on the hub, install the node, copy its key from the local UI, paste it into Settings > Link Node. The daemon now survives being started out of order instead of failing with a traceback. Adds QE/deploy/ — generic deployment (deploy-hub.sh, deploy-node.sh) kept separate from the demo scenario (demo.py, demo.env, README.md). Credentials live in QE/, which is gitignored; verified with git check-ignore. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * test: JS/Python transcript parity across the language boundaryChristophe Besson2026-08-131-0/+175
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The handshake proof and admin signature transcripts are built independently in crypto.js and in meshbay_common, and compared by producing identical bytes. Nothing on the wire carries the transcript — that is the design — but it means a one-byte disagreement between the two implementations is invisible to every other test while causing a total outage: no browser could complete a handshake with any node, and every file deletion would be rejected. Nothing else in the suite crosses this boundary. The 278 Python tests would all still pass. Drives the real crypto.js under node (stubbing window and crypto, which the module body touches but these functions do not) and compares against the real Python for the same vectors: both roles, short and empty group ids, non-ASCII group names and filenames — TextEncoder and str.encode must agree on UTF-8 — and field splits that would collide under naive concatenation. Verified to actually catch a mismatch rather than trusted for passing: removing one length prefix from the JS fails 6 vectors, and changing a single byte of the domain-separation prefix fails 6. crypto.js restored byte-identical afterwards. Skips when node is absent, which is a coverage gap rather than a pass — worth making a hard failure in CI (18.4). Tests: 168 hub+common. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * docs: draft-v5 — mark C6, M8 and 11.5.8 closedChristophe Besson2026-08-131-12/+29
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Phase 11.5 is complete. All six critical and all seven high findings from the second review are now closed, bounded, or deferred by explicit decision. Updated in place rather than appended, so the document does not carry stale "open" markers next to shipped work: - §9 split into "closed since this document was drafted" and "still open", with C6, 11.5.6, 11.5.8 and M8 moved across and the closing mechanism recorded for each - §2 claim table: node impersonation is no longer pending - §3.1 QUIC now shows the unified handshake enforced - §4.2 records what the QUIC binding actually turned out to be, including the finding that a resumed TLS session carries no certificate, so the anchor travels with the session ticket - §4.4 states that the client pins pk_node and refuses a change Added a scope note: with C6 closed, pinning is defence in depth, not the primary control. A substituted node already fails the GEK proof; pinning covers the case where an attacker holds the group key and swaps the node underneath. H3 remains the last unfixed finding, and the document still says so. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * feat(client): pin node identities on first use — closes 11.5.8Christophe Besson2026-08-132-0/+24
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | The client verified the node's Ed25519 signature but did not remember which key it had seen, so a substituted node was caught only by its lack of the GEK. Trust On First Use: the node's public key is recorded per node_id on the first successful handshake and compared on every later one. A change is refused outright — strict, per operator decision. A warning users can click through is decorative, and this is the SSH known-hosts tradeoff taken deliberately. Scope, stated honestly: with C6 closed this is defence in depth, not the primary control. A substituted node already fails the GEK proof. Pinning covers the case where an attacker HAS the group key — an ex-member, or a leaked GEK — and swaps the node underneath, which the proof alone cannot distinguish from the real one. Strict refusal needs an escape hatch or it is a dead end: a node operator who reinstalls and loses their keystore generates a new pk_node and would otherwise lock out every member. Settings gains a "Node identities" section showing the pin count and clearing them, with copy telling the user to verify out of band first. Also exposed as MeshBayTransport.clearNodePin() for the native client. Tests: hub+common green; all five static JS files syntax-checked. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * fix(hub): require proof of possession on node announce — closes M8Christophe Besson2026-08-135-13/+232
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Phase 11.5.10. POST /v1/nodes/announce accepted any pk_node with no proof the announcer held the matching private key, so a user could register a node record carrying someone else's node key, and records accumulated without limit. The announcer now signs a domain-separated message binding the key to their account — meshbay:node_announce:{user_id}:{pk_node}:{timestamp} — reusing the shape already proven by /v1/nodes/auth, so a signature for one can never satisfy the other. Same 60-second window. Re-announcing the same key now updates the existing record in place instead of creating a new row. Three test helpers had to be taught to sign, which is the useful part: nothing in the suite had ever exercised announce with an attacker's key. The new tests cover the missing proof, a foreign key, a stale timestamp, and idempotence. Note for the record: the node key is independent of the user's identity key. Two hub tests asserted the announced pk_node equalled the user's pk_ed, which happened to be true only because the daemon announces its keystore key. They now assert against the announced key itself. Tests: 157 hub+common, node suite green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * feat(quic): GEK proof and mutual authentication — closes C6Christophe Besson2026-08-133-8/+186
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Phase 11.5.4/5/6 — finding C6, the last open critical finding. QUIC ran a JWT-only handshake: a forged or stolen token reached the node and could inject chat without ever holding the group key. It now runs the same challenge/response as WebRTC through meshbay_common.handshake — client nonce, role-bound length-prefixed transcript, GEK proof, and the node proving itself with a GEK proof plus an Ed25519 signature over the transcript (C3). 11.5.6 channel binding, resolved by spike and then by two findings the spike could not predict: * aioquic 1.3.0 exposes no RFC 5705 exporter, and the peer certificate only via a private attribute. The server reads its own certificate from disk, so no internals are touched on that side; the client's access is guarded and fails loudly if an upgrade moves it. * A RESUMED TLS session carries no certificate — aioquic does not re-send it, so there is nothing live to bind to. The anchor therefore travels with the session ticket, which is sound because the ticket is cryptographically derived from the handshake where that certificate was presented. * The anchor had to travel with the ticket rather than live on the client object: resumption constructs a fresh client, so an instance-level cache was silently useless. Caught by the resumption test, not by inspection. Both paths refuse rather than degrade. No certificate and no cached anchor means the handshake fails; it never falls back to an unbound proof, which would silently drop MitM detection (L4). QuicChunkClient gains a peer_cert_der property and constructor argument, mirroring how session_ticket is already carried by the caller. Tests: 9 quic/multi-group, full node+common suite green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * docs: record 11.5.6 spike — QUIC channel binding constraintsChristophe Besson2026-08-131-1/+39
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Investigated aioquic 1.3.0 before implementing the QUIC challenge/response, since the binding anchor gates the whole design. No RFC 5705 exporter exists (aioquic.tls.Context has no export_keying_material), so the preferred anchor is unavailable. Certificate access is asymmetric: the server reaches its own cert via the public tls.certificate, but the client can only reach the server's via tls._peer_certificate — a private attribute, behind a QuicConnection that exposes no tls accessor at all. That matters because binding a security check to a private API means an upgrade can remove it silently. Since make_proof() refuses an empty binding (11.5.21), a rename would fail loudly rather than degrade — but only while the refusal path stays strict. Three options recorded with a recommendation: certificate hash via the private attribute with a guard test that fails CI on upgrade, plus pinned aioquic; or bind to pk_node instead, which for QUIC may suffice since signaling is not hub-relayed — but that requires certificate pinning, as the QUIC client currently does not verify the TLS certificate at all; or upstream an exporter. No implementation started: the challenge/response needs both protocol sides, test updates in two files and multiple verification cycles. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * docs: draft-v5 architecture specChristophe Besson2026-08-133-2/+345
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Supersedes draft-v4, which described a system the code did not implement and made several claims that were simply wrong — "ALL operations require the GEK proof" (true on one of four transports), "Argon2id 256 MB" (hub only), "hub stores no content metadata" (private file hashes were registered with it). Written as a delta over v4: sections not restated are unchanged. Carries an explicit rule — a claim must name the adversary it holds against — and a per-adversary table replacing v4's informal assurances. Records the decisions: transport (aiortc primary, QUIC retained, TCP and the node HTTP API removed), unified handshake with mutual authentication, admin operation transcripts, node authority over GEK storage and activation, upload confinement, hub node-registration and signaling authorization, and the client architecture — hub keeps serving the web SPA, native client offered alongside, hub minimization deferred. States plainly what is NOT true. The defensible claim is "the hub cannot read your content unless it actively attacks you", not "unreadable by other parties, even the hub": H3 (hub is the key directory and can substitute a key at invite time) is open until Phase 12.1, and T3 (hub serves the SPA) is accepted permanently by decision. Content is also readable by every group member and by the node operator, so "end-to-end" here means client-to-node, never client-to-client. Corrects the v4 NAT traversal account: punch_nat() is a single UDP probe with no STUN, no candidate gathering and no fallback, validated on one ISP. ICE is the traversal path, including for native clients. Open items listed with status, including C6 on the QUIC path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * docs: rework Phase 12 — hub minimization deferred by operator decisionChristophe Besson2026-08-132-48/+69
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Operator decisions (tmp-decisions.md D1/D2/D4): - the hub keeps serving the web UI (zero-install path stays) - a native desktop client is offered ALONGSIDE it, not as a replacement - hub minimization is off the critical path and may be dropped Phase 12 was "Hub minimization: registrar and nothing more". Most of it is dropped: route-inventory blindness test, opaque private-group metadata, chat_notify metadata minimization, residual schema cleanup. The swarm item already shipped in 11.5.18. Two items are kept, because the decision makes them more relevant rather than less — the hub stays in the trusted path by choice, so what it can substitute and what code it serves both still matter: 12.1 key transparency + safety numbers [H3]. This is the last open High finding and nothing else fixes it: the hub is the public key directory, so substituting a key during an invite hands it the group key silently, with no JWT forgery and no code injection. Dropping Phase 12 wholesale would have left it open indefinitely. 12.2 served-SPA integrity: CSP, SRI, and a hub-published signed digest of the bundle so a native client can verify what the browser was given. 12.3 honest labelling of /app/ as the hub-served path. 12.4 written threat model — the thing that stops the overclaiming pattern. Recorded consequence: T3 is now accepted permanently for browser users. A hub that serves the code can exfiltrate keys from the page whatever the protocol does. The claim that still holds, and that the docs should make, is "the hub cannot read your content unless it actively attacks you" — not "unreadable by other parties, even the hub". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * refactor(quic): share the unified handshake authorizationChristophe Besson2026-08-132-25/+36
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Phase 11.5.4 — QUIC half. Findings M1, M9 on this transport. quic_server._do_handshake_sync was a second, weaker copy of the WebRTC logic: group_id was optional, so omitting it skipped the membership check entirely and fell back to the node's first group (M1); node-scoped daemon tokens were accepted as client tokens (M9); and the checks could drift from the WebRTC path independently, which is how they diverged in the first place. Authorization now comes from meshbay_common.handshake, shared with WebRTC. C6 IS STILL OPEN ON THIS TRANSPORT. There is no GEK proof here yet: a forged or stolen token still reaches the node over QUIC and can inject chat without holding the group key. What remains is the challenge/response and the mutual node proof — quic_binding() is written and unit-tested for exactly this, and 11.5.6 (whether a certificate hash is the right anchor, or an RFC 5705 exporter is reachable from aioquic) is still unproven. This commit narrows the gap to the proof itself; it does not close the finding. QUIC tests updated: default tokens are members of the test group, and clients pass group_id, since it is mandatory now. Tests: 9 quic/multi-group, full node+common suite green. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * feat(mnp): unified handshake with mutual authenticationChristophe Besson2026-08-137-123/+680
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Phase 11.5.4/5/7/8 — findings C6 (WebRTC half), C3, L4, M1, M9. New meshbay_common/handshake.py is the single implementation of authorization and proof: JWT verify, scope, denylist, mandatory group_id, membership, hosting. The handshake previously existed three times over and only the newest copy enforced the GEK proof. C3 — mutual authentication. Authentication ran one way: the client proved itself, the node proved nothing. handshake_ack.node_pk was never verified against anything and per-chunk signatures had been dropped in Phase 9.15, so a peer that had hijacked signaling (C2) or been substituted by the hub could accept the client's proof, ignore it, and serve a forged index, forged chat history and a forged is_node_admin flag. The client now sends a nonce; the node answers with its own GEK proof over that nonce AND an Ed25519 signature over the transcript; the browser verifies both and refuses otherwise. It also refuses an unchallenged handshake_ack, which previously let a peer skip proving anything at all. L4 — the proof was nonce ‖ offer_fp ‖ answer_fp: bare concatenation, and a missing fingerprint silently degraded it to nonce-only, dropping MitM detection (NS5). Every field is now length-prefixed and domain-separated, the role is bound so a client proof cannot be replayed as a node proof, and an absent channel binding is refused rather than tolerated. M1 — group_id was optional; omitting it skipped the membership check entirely and fell back to the node's first group. Now mandatory. M9 — node-scoped daemon tokens are refused on the client path. NOT DONE: quic_server.py still runs its own JWT-only handshake, so C6 remains open — a forged or stolen token reaches a node over QUIC and can inject chat without holding the GEK. quic_binding() is written and unit-tested but unwired. 11.5.6 (whether the certificate-hash anchor works with aioquic, or an RFC 5705 exporter is reachable) is unproven. 11.5.8 TOFU pinning of pk_node is not done: the client verifies the node's signature but does not yet remember which key it saw last. Adds packages/meshbay-common/tests/test_handshake.py (18 tests) covering the properties every transport must inherit. WebRTC test helpers rewritten around the shared module; _make_jwt now defaults to the test group, since group_id is mandatory. Tests: 24 webrtc, 176+ node+common. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * fix: resource limits, signaling authz, node admin UI tokenChristophe Besson2026-08-137-32/+380
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Phase 11.5 — findings H6, C4 (partial), and milestone 11.5.3. H6 — resource exhaustion. Several paths let one peer degrade or stall a node: * the DataChannel frame limit was a flat 64 MB applied BEFORE authentication, so an unauthenticated peer could announce a huge frame and dribble bytes into it. Unauthenticated peers now get 64 KB; the large budget is granted only after the GEK proof, where it is needed for uploads. * _do_stream_segment ran subprocess.run(..., timeout=30) directly in the event loop, stalling the entire daemon — every peer, every group — for up to thirty seconds per request. Now async, with a timeout and process kill. * ffmpeg was spawned per stream request with no cap. Both streaming paths now share a transport-wide semaphore. * POST /v1/nodes/{id}/webrtc/offer was reachable by any authenticated user for any node, with no membership check and no rate limit, making the target node allocate an aiortc PeerConnection and gather ICE on demand — remote resource exhaustion against a third party's machine. Now rate limited, capped per user, SDP size bounded, and the caller must share an active group with the node. That also closes the H4 gap where signaling ignored group status. * POST /v1/nodes/{id}/incoming took peer_ip verbatim, so any user could make an arbitrary node emit UDP packets to an address of their choosing. The probe target must now match the caller's own source address. C4 (partial) — the pre-proof bundle window. GEK and keypair bundle fetches are served before the GEK proof by necessity: the client needs its wrapped bundle in order to compute the proof. That window is a disclosure surface a hub can reach by forging a JWT. Bounded to 4 fetches per session and audited as "pre_proof_fetch". The real fix is removing remote keypair bundles entirely, which belongs to the native client (Phase 13.3). 11.5.3 — the node admin UI was unauthenticated because it binds loopback. But any local process can reach it, and so can a page in the operator's browser via DNS rebinding — and this API re-initialises group keys and reads the audit log. H2 showed script execution there equals full control. Now gated by a per-run token, printed at startup, accepted as ?t= or X-MeshBay-Token. One test needed rewriting rather than adding: the first version asserted "subprocess.run(" was absent from the source, which also matched the comment documenting the old behaviour. It now parses the AST and checks the property. Tests: 121 node, 142 hub+common. Regression suite 47 node + 10 hub. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * fix: swarm privacy, revocation persistence, keystore KDF, audit integrityChristophe Besson2026-08-1314-66/+406
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Phase 11.5 hardening batch — H7, H4, M2, M6, M7, L1, L3, L6. H7 — private content hashes leaked to the hub. The daemon registered blake3 hashes for every group it hosted, private ones included, giving the hub a content fingerprint of every private file and letting anyone confirm whether a known file exists in the network. The leak was dormant only because the routes were declared on the groups router with a full path and mounted at /v1/groups/v1/swarm/* — the node's calls 404'd into a swallowed exception. Fixing the path alone would have activated the leak, so both land together: registration is gated on group visibility, the routes moved to a real /v1/swarm router, and the lookup now requires authentication. H4 — revocation was advisory. Group revocations were signed and broadcast by the hub and then dropped by the node, whose handler understood only "user" and "jti", so "suspend a group" enforced nothing. The denylist was also in-memory only, so a restart silently un-revoked everyone. Now persisted to data_dir/denylist.json, group targets honoured on both transports, and live sessions for a revoked group are closed. M2 — the node keystore, which protects the node's Ed25519 and X25519 private keys, was still deriving at 64 MB long after the hub's password verifier moved to 256 MB; the docs recorded the bump as done, true for the hub only. Raising the constant alone would have made every existing keystore permanently undecryptable, so envelopes now record the parameters they were written with and pre-M2 files continue to open under the legacy profile. M6 — registration inserted its audit row with a NULL user_id and then ran UPDATE ip_logs SET user_id=<new> WHERE user_id IS NULL, claiming every unattributed row in the table: failed logins for other usernames, concurrent registrations. In logs retained a year for legal requests, that attributed other people's connections to the wrong account. M7 — X-Forwarded-For was trusted unconditionally at four call sites, so anyone could forge the IP written to the compliance log and evade per-IP rate limits. New netutil.client_ip honours the header only from a trusted proxy and takes the rightmost hop (the one our proxy appended); no direct header reads remain. L1 dead GEK_REQUEST/GEK_RESPONSE constants removed; L3 peer errors no longer echo exception text (paths, internal state); L6 email sanity-checked instead of accepting any string — deliberately not RFC 5322, to avoid a new dependency. test_daemon_index_change_pushes_to_peers asserted that a PRIVATE group's hashes are registered with the hub. Split: private asserts not-called (index push to members still asserted), and a new test proves public groups still register. That is the fourth pre-existing test found asserting a vulnerability as intended behaviour, after gek auto-activation, the transport-wide chat_store and the blind admin challenge. Tests: 116 node, 132 hub+common. Regression suite now 43. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * fix(hub): authenticate node WebSocket registrationChristophe Besson2026-08-132-13/+256
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Phase 11.5 — finding C2 (see second-review.md). /v1/nodes/ws took node_id and group_ids straight from the client's first message with no ownership check: node_id = msg.get("node_id") or decoded.get("sub", "unknown") _connected_nodes[node_id] = ws Any registered user could connect with an ordinary browser token, claim a victim node's id and overwrite its entry. Every WebRTC offer for that node was then relayed to the attacker, who answered with their own SDP — full node impersonation. The DTLS channel binding does not help, because the attacker is the endpoint rather than a relay: the browser sends its GEK proof to the attacker, who ignores it and replies handshake_ack. The attacker received the victim's encrypted keypair bundle, chat and uploads, and could serve a forged index. Registration now requires scope == "node", verifies Node.user_id against the token subject, checks the account is active, and refuses to displace a live registration instead of silently overwriting it. group_ids are intersected with the operator's actual membership: a node may narrow the set to what it hosts but cannot widen it, so it cannot advertise itself as an online source for arbitrary groups. Authorization uses a short-lived session rather than Depends(get_db): a node WebSocket lives for hours and a request-scoped dependency would pin a PostgreSQL connection for its whole lifetime. BEHAVIOUR: a node hosting a group whose hub membership was never recorded for the operator's account will stop appearing in GET /v1/groups/{id}/nodes. Adds tests/test_node_ws_auth.py (7 tests). The node WebSocket had no test coverage at all, which is why this went unnoticed. Tests: 109 node, 139 hub+common. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * fix(node): group isolation, upload confinement, GEK seizure, admin challengeChristophe Besson2026-08-1311-155/+919
| | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Phase 11.5 — findings H1, C5a, H2, C5b, H5 (see second-review.md). Batched together because the node-side changes share webrtc_server.py and cannot be separated into working commits. H1 — cross-group chat leak. chat_store, the peer registry and the display-name cache were read from the shared transport context, and daemon.py hoisted the FIRST group's chat store onto it. On a node hosting several groups every group's messages went to one database, chat_history served them back to members of every other group, and chat broadcast reached all peers regardless of group. All three now resolve through _group_ctx(). C5a — upload confinement. Uploads landed in the shared root under a client-chosen name and overwrote whatever was there. Any member could destroy the operator's files, and by becoming the recorded uploader of the replaced file could then delete it through the uploader path, bypassing the Ed25519 admin challenge. Uploads now go to a per-user quarantine (.uploads/{user_id}/), refuse to overwrite, and enforce chunk ordering, a filename allowlist and a size cap. H2 — stored XSS in the node admin UI. Filenames chosen by any group member were interpolated raw into the localhost UI, which has no authentication, so script execution there equals control of the node admin API. Now html.escape() throughout, textContent in the audit table, plus CSP/nosniff/no-referrer. The CSP contains exfiltration but cannot stop injected inline script — escaping is the fix. C5b — group key seizure. gek_bundle_store wrote whatever any member sent and auto-activated bundles addressed to the node operator. The operator's X25519 public key is public (the node publishes it in handshake_ack), so any member could wrap a key of their choosing for it and take over the group, locking every legitimate member out. Storing now requires an operator signature and _try_activate_gek is removed: nothing arriving over MNP can set a live GEK. H5 — unbound signing oracle. The node challenged with 32 raw random bytes and the client signed them blind, so a signature named no operation, subject, node or time. New meshbay_common/adminop.py defines a length-prefixed, domain-separated transcript; both sides build it independently and the client refuses to sign when the announced op/subject do not match its request. BREAKING: a group admin who does not operate the node can no longer store GEK bundles on it. Invites must be performed by the node operator. Adds tests/test_security_regressions.py. Verified against pre-fix source via git stash. Three pre-existing tests asserted the vulnerable behaviour as a feature and were inverted: gek auto-activation, and the transport-wide chat_store in test_daemon. Tests: 109 node, 132 hub+common. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| * fix(node)!: remove unauthenticated HTTP file API and TCP transportChristophe Besson2026-08-1312-1346/+45
|/ | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | | Phase 11.5.A — findings C1 and C6 (see second-review.md). C1: the per-group HTTP file API bound 0.0.0.0 for every configured group, private ones included, and served two endpoints with no authentication at all: GET /index (full Mesh Group Index) and GET /file/{id} (raw plaintext file via FileResponse). Anyone able to reach the port — LAN, forwarded port, permissive IPv6 — read every private file. This bypassed the entire GEK-proof and node sovereignty layer. Deleted rather than patched: it duplicated MNP without any of its controls. C6: the TCP+TLS chunk server accepted a bare JWT with no GEK proof, leaving a second non-compliant handshake path. Deleted; QUIC remains and will be brought to parity with WebRTC by the unified handshake in 11.5.4. Transport decision recorded in transport/__init__.py: WebRTC/ICE is primary for browser and native clients (the only NAT traversal validated here — 2 ISPs, IPv4 STUN + IPv6, 4G CGNAT); QUIC is kept for LAN, port-forwarded and hub-less group:// access. punch_nat() is a direct-connection helper, not a traversal stack. Also removed server_ssl_context()/client_ssl_context() from tls_cert.py (no remaining callers) and a dead import of the former in quic_server.py. generate_self_signed_cert() stays: QUIC uses it, and the certificate hash is the intended channel-binding anchor for 11.5.6, since QUIC has no DTLS fingerprint to bind the GEK proof to. BREAKING CHANGE: node.toml keys `port` and `http_port` are gone. Regenerate config with `meshbay-node init`. Env var MESHBAY_PORT -> MESHBAY_QUIC_PORT. Tests: 198 passed (209 - 7 test_http_server - 4 test_transport). No other test changed status. Net -1300 lines. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>