diff options
Diffstat (limited to 'docs/transfers-v1.md')
| -rw-r--r-- | docs/transfers-v1.md | 1159 |
1 files changed, 1159 insertions, 0 deletions
diff --git a/docs/transfers-v1.md b/docs/transfers-v1.md new file mode 100644 index 0000000..f99424f --- /dev/null +++ b/docs/transfers-v1.md @@ -0,0 +1,1159 @@ +# Transfer slots, queueing, pause and resume — build record + +> **Status: built.** Leases, caps, queueing, the settings surface, the client +> widget, pause and resume all shipped; the compulsory-lease flag day is +> **MNP 3.0**. +> +> **The design is `MESHBAY_DESIGN.md` §5.5**, which states the lease model, the +> caps, and the reasoning behind the browsing exemption. Read that first. This +> document is kept for what a synthesis cannot carry: the failure-mode analysis +> in §5 (every way a slot can be lost, every way a client can be left waiting, +> and the invariant that checks both), the target-by-target resumability table in +> §6, and §12 — what a live pass found after the work was called done, which is +> the most useful part of the file. +> +> **Where this document and the code disagree, the code is authoritative**, and +> two places are known to disagree: the leaseless-read bound is 12 files per +> session with an idle expiry, not the 2 written in §3.4.1, and the wire section +> describes MNP 3.0 as pending when it has shipped. `MESHBAY_DESIGN.md` §5.5 +> carries the current values. +> +> Convention carried over and not negotiable: **a claim here names the failure it +> holds against.** "A slot is released when the transfer finishes" is not a +> design; "here is every way a slot can be lost and what reclaims it in each +> case" is. + +--- + +## 0. The short answer + +| Asked for | Feasible | Where the work is | +|---|---|---| +| `max_concurrent_downloads` / `max_concurrent_uploads` per node (default 8) | Yes | The node has **no concept of a transfer** today — that is the whole job | +| Per-member caps inside a group (default 2) | Yes | New per-group operator-signed setting, same shape as `apps_enabled` | +| "Max reached" refusal + client queues the rest | Yes | Better as a **node-side queue with an explicit grant**, not a client retry loop | +| A visible **waiting** state in the transfers widget | Yes | `TransferStore` already has a status field; it gains three states | +| Cancel | **Already works** (`transfers.cancel`, `transfers.js`) | Needs to also tell the node, which today it does not | +| Browsing — posters, thumbnails, previews, listings — untouched by the caps | Yes | §3.4. Structural, not a threshold: the three functions that draw a widget row are the three that take a slot | +| Pause with resume | Yes, **in three tiers** — see §6. Not every download target can be resumed, and the interface must not offer what a target cannot do | + +The risk is not in any one of these. It is that a capped resource plus a queue +is exactly the shape of bug that shows up as *"my download says waiting and +never starts"*, days later, on someone else's machine. §5 is therefore the +longest section in this document and is the one to review hardest. + +--- + +## 1. What exists today, measured + +Read this before designing anything: three of the assumptions one would +naturally make are false here. + +**A download is invisible to the node.** `pipelinedDownload` +(`static/file-utils.js`) sends `PIPELINE_WINDOW = 8` independent `file_req` +messages and reassembles the answers. Each one is served by +`_do_file_request` (`webrtc_server.py:3380`), which looks the entry up, reads +and encrypts one 1 MB chunk, waits for room on the channel +(`DOWNLOAD_BUFFER_HIGH`, `webrtc_server.py:253`) and sends it. **The node never +learns that a download started, and never learns that one ended.** There is +nothing to count and nothing to cap. This is the central fact of this work. + +**An upload is half-visible.** `_do_file_upload` (`webrtc_server.py:4471`) +keeps `self._uploads[f"{rel_dir}/{filename}"]` with `next_index` and a `.part` +file on disk. That state is **per session object and in memory**: a browser +that disconnects mid-upload leaves a `.part` file on the operator's disk that +nothing ever removes, and a retry starts at chunk 0 under a new `_free_name`. +So there is already a resume story here, it is just not reachable — and there +is already a leak. + +**Video streaming is capped, and its cap is a good model.** `_stream_video` +(`webrtc_server.py:5230`) takes `ctx["_transcode_sem"]`, refuses with +`"Server busy, retry shortly"` when it is empty, and every hard-won lesson in +this repo about slots — `_replace_stream`, `shutdown_tasks()`, `_spawn()` and +the garbage-collected task, `await proc.wait()` after `kill()` — is a lesson +about *not losing a slot*. That history is why §5 exists. + +**One bug found while reading, in scope for this work.** `ops.py:1362-1365` +hot-swaps the stream cap by assigning `webrtc._stream_sem` — an attribute that +does not exist. The real semaphore is `ctx["_transcode_sem"]` +(`webrtc_server.py:5220`), and `hasattr(webrtc, '_stream_sem')` is always +False, so **changing `max_concurrent_streams` from the Node page has never +taken effect without a restart**, contrary to §2.11 of draft v6. Two more +hot-swappable caps are about to be added next to it; fix it with one shared +helper rather than three copies of the same mistake. + +**Existing client state.** `TransferStore` (`static/transfers.js`) already +survives leaving a group, already holds the transport open until the last +transfer finishes (`releaseWhenIdle`), already cancels on sign-out, and already +computes a windowed speed. Its statuses are `running | done | cancelled | +failed` and its `run()` contract is a promise plus a polled `signal.aborted`. +That contract is the thing that has to change, and it is the only client-side +structural change. + +--- + +## 2. What a naive implementation gets wrong + +Worth stating, because each of these is a plausible first attempt. + +1. **Counting `file_req` messages.** A rate limit on chunk requests caps + throughput, not concurrency, and gives the client nothing to render as + "waiting". It also cannot distinguish a download from a thumbnail. +2. **Inferring a transfer from activity.** "A (session, file) pair with a chunk + request in the last 30 s is a live download" needs no protocol change and is + tempting. It makes the admission decision unobservable — the client cannot + be told *why* it is being refused, cannot be told when to try again, and a + paused transfer is indistinguishable from a crashed one. +3. **A client-side queue only.** The client already knows what it wants to + download; it does not know what the other five members are doing. A queue + held only by clients cannot order anything fairly and turns into a retry + storm against a busy node. +4. **A refusal the client retries on a timer.** This is the version that + produces the "stuck at waiting" report: the client polls, the node refuses, + nobody is at fault, and there is no evidence anywhere. The node must + **push** the grant. +5. **Gating every `file_req`.** Thumbnails, posters, cover art and audio + transcodes are served through the *same* `file_req` path + (`_try_serve_thumbnail`, `webrtc_server.py:3353`), and so are image and text + previews. Requiring a slot for all of them means opening a poster grid + queues behind a film. §4.4 handles this. +6. **Holding a slot while paused.** If pausing keeps the slot, one member can + pause eight downloads and close the node to everyone else. Pausing must + release. + +--- + +## 3. The design + +### 3.1 The lease + +Introduce **the transfer lease**: the node's record that a peer is transferring +something, held for the length of the transfer and released by name. + +``` +transfer_open { tr, kind: "download"|"upload", bytes, chunks, label_hint } + → transfer_state { tr, state: "granted" } + → transfer_state { tr, state: "queued", ahead: 3 } + …later… + → transfer_state { tr, state: "granted" } ← pushed, not polled +transfer_close { tr, reason: "done"|"cancelled"|"paused"|"failed" } + → transfer_state { tr, state: "closed" } +``` + +Six properties, each of which is load-bearing: + +- **`tr` is drawn by the client**, 16 random bytes hex, exactly like + `upload_id` (`transport.js:2208`). Re-opening after a reconnect with the same + `tr` is idempotent, so a reconnect cannot double-charge a member for one + transfer. +- **A lease is scoped to the connection, never to the account.** It dies with + the session, which is what makes the primary reclaim deterministic (§5). +- **A lease covers a job, not a file.** A directory zip + (`downloadDirectory`, `file-utils.js`) is dozens of files and is **one** + lease. So is a resumed transfer. +- **A lease is never persisted.** Not in `roster.db`, not on disk, not across a + node restart. A restart drops every session anyway; a lease that outlived the + process would be a slot nothing can release. +- **The node counts leases, not bytes.** What a slot protects is concurrency — + open file handles, disk seeks, and the channel buffer each transfer keeps + full — not bandwidth. +- **Every `file_req` for a real file carries its `tr`.** The exceptions are in + §4.4 and they are the delicate part of this whole design. + +### 3.2 Two caps, in one order + +| Cap | Scope | Default | Where it lives | +|---|---|---|---| +| `max_concurrent_downloads` | node-wide, all groups | 8 | `[node]` in `node.toml`, override in `roster.db` (§7) | +| `max_concurrent_uploads` | node-wide, all groups | 8 | idem | +| `max_downloads_per_member` | one group, one account, **across all their devices** | 2 | per-group, operator-signed (§7) | +| `max_uploads_per_member` | one group, one account, across devices | 2 | idem | + +Downloads and uploads have **separate pools**, and video streaming keeps its +own third pool (`max_concurrent_streams`) untouched. A member watching a film +is not charged a download slot, and a download does not make the next film +answer "server busy" — those are different resources with different costs and +merging them would make both caps meaningless. + +**Order of checks, and it matters:** per-member first, then node-wide. A member +at their own cap queues *behind their own transfers* and never consumes a +node-wide slot they would then hold while a second member has none. Reversed, +one member arriving first takes all eight. + +**"Per member" means per account, summed across their devices**, resolved with +`_sessions_of(user_id)` (`webrtc_server.py:3325`) — which exists for exactly +this reason, since device linking landed. Two browsers and a desktop client +signed in as the same person share the two slots. Anything else makes the cap a +function of how many tabs someone opens. + +### 3.3 The queue + +One FIFO per pool, per node, holding `(tr, session, user_id, group_id, +enqueued_at)`. + +- **Drained in one place.** `_release_slot()` is the only function that returns + a slot, and it is the only caller of `_pump_queue()`. Every path that ends a + transfer goes through it, in a `finally`. Two functions that both release + would be this repo's flow-control-accounting lesson (`CLAUDE.md`) one feature + later. +- **Head-of-line blocking is skipped, not waited on.** When a node-wide slot + frees, walk the queue and grant to the first entry whose *member* is under + their own cap. Granting strictly in order would let one member at their + personal cap stall the whole node. +- **A grant has an acceptance deadline.** 30 s. If no `file_req` bearing that + `tr` arrives, the grant is revoked (`state: "queued"` again, at the tail) and + passed on. Without this, a client that dies between the grant and the first + chunk holds a slot until the idle timeout. +- **The queue is bounded**: 32 entries per account per pool. Beyond that the + node answers `too_many_queued` and the client holds the rest in its own list, + sending `transfer_open` as its own transfers finish. Unbounded queues are how + a node runs out of memory politely. +- **Positions are pushed, throttled.** After each pump, recompute positions and + send `transfer_state {state:"queued", ahead:n}` only to the sessions whose + number changed, at most once every 2 s per session. + +### 3.4 What is *not* gated — posters, thumbnails, previews, navigation + +**Operator decision, 2026-09-08: browsing a group is never subject to a +transfer slot.** Not the poster grid, not the album covers, not the video +thumbnails, not the file list, not opening a photo or a PDF to look at it. A +member must be able to browse a group that is at capacity exactly as they +browse an idle one. This is a requirement, not a tuning parameter, and the +sections below are written to satisfy it structurally rather than by choosing a +lucky threshold. + +**Navigation proper never touches this path at all.** The file list, the +directory tree, the poster metadata and the album metadata travel as +`index_sync`, `index_delta`, `media_meta_req`, `music_meta_req` and +`link_preview_req` — different message types, sealed under the group key, with +no relationship to `file_req`. Nothing in this design can reach them. That half +needs no rule; it needs only to be verified by a test that fails if someone +later routes a listing through `file_req`. + +The `file_req` path carries three genuinely different things, and they are +distinguishable **structurally**, by what the id resolves to and by which +function asked: + +| What | Call site | Resolves to | Rule | +|---|---|---|---| +| Thumbnails, TMDB posters, cover art, cached audio transcodes | `MediaThumb` (`video-app.js:204`), and the Music/Photos grids through it | a **media-cache id**, not an index entry — `_try_serve_thumbnail`, `webrtc_server.py:3353` | **Never leased, never counted, never queued.** One chunk each, out of a bounded cache the node built itself | +| Looking at one file — a photo opened full size, a PDF, an image, a text file | `PhotoViewer` (`photos-app.js:171`), the Files preview modal (`files-app.js:610`) | a real index entry, fetched whole | **Not leased.** Bounded by §3.4.1 below, which no real viewer ever reaches | +| Downloading, and uploading | `downloadEntry`, `downloadDirectory` (`file-utils.js`), `uploadFile` (`transport.js:2192`) | a real index entry | **Leased.** These are exactly the three call sites that go through `transfers.start()` — the three that produce a row in the transfers widget | + +The last column is the whole rule, and it is worth stating as a sentence +someone can check by reading: **a transfer is something the transfers widget +shows. If it does not appear in that panel, it does not take a slot.** The two +sets are the same three functions, which is what makes this verifiable rather +than a matter of judgement at each new call site. + +#### 3.4.1 The bound on leaseless reads + +A `file_req` with no `tr` on a real index entry is served, subject to one +limit: **at most 2 distinct file ids in flight leaselessly per session**, with +no queue — the third is refused with `transfer_required`. + +Why this shape, and not the byte budget an earlier draft of this document +proposed: + +- **A viewer looks at one file.** The photo viewer shows one photo, the preview + modal one document. Two is already one more than any of them needs, and is + there so that prefetching the next photo stays possible. +- **A size threshold does not work here.** A RAW photo out of a camera is + 60–80 MB and is *browsing*; a 40 MB archive is a *download*. Size does not + separate them, and any threshold that let the photo through would let the + archive through too. What separates them is which function asked. +- **A byte-rate budget does not work either.** It would have to be large enough + for that same RAW photo, at which point it is large enough to be a download + channel. Concurrency is the thing being rationed, so concurrency is what the + exemption is expressed in. +- **What it costs.** A client that lies — labelling a bulk download as a view — + gets two files at a time per session instead of its member cap. That is the + residual, it is bounded, it is audited (`file_download` already goes to + `audit.db` on chunk 0), and it is the same class of statement as the cap + itself: **this is a fairness control among cooperating clients**, in the + company of `max_concurrent_streams`. It is not a defence against a member + determined to saturate the node's disk, and must never be described as one — + that member is a member, and the answer to them is `member revoke`. + +#### 3.4.2 The consequence for the interface + +A preview never shows "waiting", because a preview never queues. If the +2-in-flight bound is somehow reached, the modal reports a plain error and the +person tries again — it does not silently become a queued transfer in a panel +they were not looking at. + +### 3.5 Where the numbers are visible + +The node answers `transfer_state` with the current picture — `used`, `cap`, +`ahead` — so the client can say *"waiting — 2 of 2 of your slots are busy"* +rather than a bare spinner. The same counters go to: + +- the loopback API, `GET /api/transfers`, so the operator can see live leases + and queue depth from the CLI and the Node page, and +- a DEBUG line every 30 s: `transfer: d=3/8 u=1/8 queued=2 (skipped=1)`. + +That line is the `client_diag` lesson applied here: when someone reports a +transfer stuck at "waiting", this is the only thing that will say whether the +node ever had them in a queue. + +--- + +## 4. Wire protocol — MNP 3.0 + +### 4.1 New message types (`meshbay_common/protocol.py`) + +| Type | Direction | Carries | +|---|---|---| +| `transfer_open` | client → node | `tr`, `kind`, `bytes`, `chunks`, `from_chunk` | +| `transfer_close` | client → node | `tr`, `reason` | +| `transfer_state` | node → client | `tr`, `state`, `ahead`, `used`, `cap`, `scope` | + +`transfer_state` is the only reply, for granted, queued, revoked and closed +alike. One message type with a state field, rather than four types, because a +client that must switch on the type to find out it is still waiting is a client +that will get one branch wrong. + +`file_req` gains an optional `tr`. `file_upload` gains `tr` beside the +`upload_id` that is already there in clear. + +**Sealing.** `transfer_open` / `transfer_close` / `transfer_state` carry no +file names and no paths — `tr` is opaque, `bytes` and `chunks` are numbers — +so they stay in clear like `index_progress`, and for the same stated reason. +`label_hint` is **not** in the protocol: the client already knows what it named +the transfer, and putting a filename on the wire in clear to make a log line +prettier is exactly the trade `groupbox.py` exists to refuse. + +### 4.2 Why this is MNP 3.0 and not an additive 2.x + +The messages are additive; **the requirement is not**. A 2.0 client downloading +a 4 GB film sends no `tr`, is treated as a leaseless read, and is refused as +soon as it opens a third file (§3.4.1) — or, worse, is *not* refused and +transfers outside every cap. An opt-in switch — "enforce leases only for +clients that speak 3.0" — leaves a leaseless branch reachable on every node, +which is finding C6's lesson (`TCP accepted a bare JWT`) one feature later, and +it was already refused once for chat encryption on 2026-09-07. + +So: **`MNP_VERSION = "3.0"`, `MNP_MIN_SUPPORTED = "3.0"`**, and a 2.x peer is +refused at the handshake with `version_too_old` rather than admitted and then +mysteriously unable to download. Same play, same reasoning, and +`handshake.py`'s version range is the mechanism that already exists to pay for +it. + +**What that costs, stated plainly.** The SPA deploys with the hub, so browsers +get the new client. The **desktop client ships its own UI**, so an un-updated +app is locked out until its user updates — the version-skew consequence draft +v6 flagged under "Shipping the UI in a package". Before this ships, +`GET /v1/hub/version` needs its minimum-client-version field populated and the +client needs to show a real message ("this version can no longer connect, +update here") rather than a handshake refusal. + +### 4.3 Handshake ack + +Two fields inside the sealed configuration block (`_complete_handshake`, +`webrtc_server.py:852`), so the interface can draw correct numbers before +anything is transferred: + +```python +"transfer_limits": {"downloads": 2, "uploads": 2}, # this member, this group +``` + +Absent reads as "no limit known" and the client simply does not draw the +"1 of 2" hint — never as "unlimited", which would have the interface +contradicting the node. + +--- + +## 5. Nothing gets stuck + +The requirement the operator stated first, and the one worth over-engineering. +Two directions, and both must be closed: + +- **a slot the node never gets back** — the node fills up and everyone queues + forever; +- **a transfer the client shows as waiting when the node has forgotten it** — + one person's widget lies while the node is idle. + +### 5.1 Every way a slot can be lost + +| How it ends | What reclaims the slot | When | +|---|---|---| +| Transfer completes | Client sends `transfer_close{done}` | Immediately | +| …and the client forgets to | **Node auto-closes on serving the last chunk** — it knows `entry.size` and `CHUNK_SIZE`, so it knows the final index | Immediately | +| User cancels | `transfer_close{cancelled}` | Immediately | +| User pauses | `transfer_close{paused}` | Immediately | +| Tab closed, browser quit, app killed | `shutdown_tasks()` / `_unregister_peer()` releases every lease of the session | On WebRTC `connectionstatechange` — the same hook that already ends streams | +| Network drops, no clean close | Idle reclaim: no `file_req` under this `tr` for 120 s | ≤ 120 s | +| Client granted a slot and never used it | Acceptance deadline | 30 s | +| …and it *is* using it, but nothing said so | **`touch()` on every `file_req` carrying the lease.** Missing from the first build: the pool had the method, the tests covered it, and no caller existed — so `used` stayed False for every download and the acceptance deadline revoked each grant at 30 s while the file transferred at 20 MB/s | — | +| A grant revoked, requeued, granted again, revoked again | **Bounded at three misses, then the lease is closed.** Also missing at first: the requeue was a permanent cycle and the node logged the same reclaim every 30 s until it restarted | — | +| Client crashes mid-zip between two files | Idle reclaim (the gap between files is milliseconds) | ≤ 120 s | +| Node restarts | Leases are in memory only | Immediately | +| Group detached / member revoked mid-transfer | The existing revocation path drops the sessions, which releases their leases | Immediately | + +The first two rows are the answer to *"a user must not stay blocked once their +downloads have been done"*: the client says so, **and** the node concludes it +independently. Neither is trusted alone. + +### 5.2 Every way a client can be left waiting + +| Failure | What fixes it | +|---|---| +| Node granted a slot, the push was lost | Client watchdog: no state change for 60 s while `queued` → re-send `transfer_open` (idempotent on `tr`) | +| Reconnect: session gone, leases gone | `_onReconnected` (already in `transport.js`) re-opens the lease for every running/queued transfer, with `from_chunk` | +| Node forgot the lease (idle reclaim during a stall) | Node pushes `transfer_state{state:"revoked"}`; the widget shows **interrupted — resume**, not a silent hang | +| Node at cap and everyone is idle | Cannot happen if §5.1 holds; if it does, `GET /api/transfers` shows the leases and the operator can force-release one. Ship that endpoint | +| Queue entry for a member who left the group | Pump skips entries whose session is closed and drops them | + +### 5.3 The invariant, and how it is checked + +> **A new grant is never made past a cap, and every queue entry names a lease +> that exists. A lease is created in exactly one function and destroyed in +> exactly one function.** +> +> Not "granted leases never exceed the cap" — that was the first wording and the +> property test rejected it within a second of being written. Lowering a cap +> never interrupts a running transfer, so the count legitimately sits above the +> new value until those finish. The invariant is about what may be *handed out*, +> not about what is held. + +Two tests, not one: + +- a unit test per row of both tables above; +- a **randomised property test** — a few thousand random sequences of open, + close, drop, reconnect, pause, resume and cap changes, asserting after each + step that the counter equals the number of live leases and that no queue + entry references a dead session. The stuck-slot bug is a race by nature, and + "it works now" is not evidence against a race (`window_leak.mjs` is the + precedent in this repo). + +--- + +## 6. Pause and resume + +Cancelling already works client-side (`transfers.cancel`) and needs only to +send `transfer_close`. Pausing is new, and **whether it can be resumed depends +entirely on where the file is being written**. The interface must offer only +what the target can actually do — a pause button that quietly restarts the +download from zero is worse than no pause button. + +### 6.1 By download target + +| Target | Platform | Pause in-session | Resume after reload | Notes | +|---|---|---|---|---| +| File System Access (`showSaveFilePicker` / granted folder) | Chrome, Edge | Yes — keep the `writable` open, stop asking for chunks | **Yes** — reopen with `createWritable({keepExistingData:true})` and `write({type:"write", position})` | The full story. `downloads.js` already keeps the directory handle in IndexedDB | +| Electron native sink | Desktop app | Yes | **Yes**, once `main.js` gains a `save:resume` opening the file with `flags:'a'` and returning its current size | ~30 lines in `main.js` + `preload.js` (`main.js:797`, `preload.js:150`) | +| Service-worker stream | Firefox, Safari | **No** | No | Not on this target — see §6.5, which says what would be needed to change that. **Hide the pause button**, keep cancel | +| Blob in memory | Fallback | Yes (it is just an array) | No | Bounded by `BLOB_LIMIT` anyway | +| **Upload** (any platform) | All | **Yes** | **Yes**, via the sealed probe | A `File` is seekable and the *node* keeps the position, so there is no local target to consult. Missed entirely when 7a shipped — pause was built around the download path and uploads were refused it by the same guard that protects a transfer which cannot re-acquire its slot | + +**Corrected while implementing 7a.** `platform.capabilities` is the wrong home +for this: resumability is a property of the *target*, not of the platform. The +same Chrome yields a pausable target from a granted folder and an unpausable one +from the service worker, on the same page, for two files in the same batch. So +each target declares `pausable` itself, `prepare` carries it into the store, and +the widget renders from that record. The precedent to +follow — and the mistake not to repeat — is `_openDownloadTarget`'s silent +fallback chain (`CLAUDE.md`: *"a fallback chain reaches its floor silently"*). +Whatever the target ends up being, the transfer records which tier it got, and +the widget renders from that record. + +### 6.2 Pause releases the slot + +Stated again because it is a design decision, not an implementation detail: a +paused transfer holds nothing. Resuming rejoins the queue at the tail, and the +widget says so (*"resuming — 2 ahead"*). Anything else lets one member close +the node by pausing. + +### 6.3 Resuming the transfer itself + +Trivially available, because `pipelinedDownload` is already indexed by chunk: +resume is `nextRecv = Math.floor(bytesWritten / CHUNK_SIZE)` with the window +refired from there. Two constraints: + +- **`bytesWritten` must be a multiple of `CHUNK_SIZE`.** It is, because writes + are sequential and whole chunks — but a resumed file whose size is not a + chunk multiple (a partial write interrupted by a crash) must be **truncated + down** to the last whole chunk before resuming, never appended to. A silently + corrupted download is worse than a failed one. +- **The file must not have changed.** The entry id is its blake3 + (`GroupIndex`), so the resume record stores `{fileId, groupId, size, + bytesWritten, targetRef}`. A resume whose `fileId` is no longer in the index + fails with "this file has changed on the node", which is the truth. + +### 6.4 Resuming an upload + +The node already has the state (`self._uploads`, `.part` on disk). Make it +usable and stop it leaking: + +- **Corrected while implementing.** `transfer_open` travels in clear, and the + node identifies an upload by (member, directory, filename) — so asking there + would put the operator's filenames on an unsealed message, which is precisely + what sealing this path bought in MNP 2.0. The question is asked inside the + seal instead, as an ordinary `file_upload` with no bytes and + `UPLOAD_PROBE_INDEX` (-1); the node answers `resume_from` in the sealed ack, + writing nothing and reserving nothing. A node that predates it refuses the + index, which reads as "start from the beginning". +- Move `self._uploads` from the session to the **group context**, keyed by + `(user_id, rel_dir, filename)`, so a reconnect finds it. It is already + authorised per member; the session was never the right owner. +- **Reap orphans.** A `.part` older than 24 h with no live lease is deleted, on + a timer and at startup. This is a pre-existing leak (§1) and this work is + where it gets fixed. +- The no-overwrite rule, the filename allowlist, the size cap and the chunk + ordering are **unchanged** — draft v6 §2.1 names those four as what makes an + upload safe, and none of them is touched here. + +--- + +### 6.5 Firefox and Safari: OPFS, and what a pause would cost + +> **Corrected three times on 2026-09-08, ending here.** This section first said pause +> was impossible on Firefox; then that OPFS was the missing streaming target. +> **Both were wrong, and measurement is what settled it.** Firefox 154's OPFS +> quota is exactly 10% of the volume's size — ~12.8 GB on a 128 GB disk, hit +> *mid-download* — so OPFS cannot carry a film. The unbounded path is the +> **service worker**, which was already there and merely unreliable; it was +> fixed on 2026-09-08 (see `memory-audit-large-files.md` §6.5 and +> `test_streamed_download_reliability.py`) — and then fixed again the same +> evening, because "reliable" turned out not to include *staying alive*: an idle +> service worker is terminated after about thirty seconds and a streaming +> response does not count as activity, so every download longer than that lost +> its reader mid-file (§12.2). What survives below is only the pause question, +> for which OPFS is still the candidate, under its quota. + +Deferred by decision, so this section records **why**, because "Firefox cannot +pause" is not true as stated and would be a bad thing to leave in a plan. + +**Why the current target cannot pause.** On the service-worker path +(`downloads.js:openStreamedDownload`) the browser is already writing an HTTP +response to its own download folder. Three separate things break, and only the +first is about our code: + +- pausing means not writing to the stream, which leaves the browser's download + stalled in its own download manager — visible to the user, outside our + control, and cancellable by them in a way we never hear about; +- **a service worker is terminated when idle.** Firefox and Chromium both kill + one after tens of seconds with nothing to do. A pause that outlives that — + which any pause waiting on a queue does — takes the stream down with it; +- resume after a reload is genuinely impossible on this target. The response is + committed; no API reopens a browser-managed download to append to it. + +**The alternative that does exist: OPFS.** The Origin Private File System +(`navigator.storage.getDirectory()`) is a real, seekable, persistent +filesystem, private to the origin, and it **is** implemented in Firefox and +Safari — which is exactly what those two lack when they lack File System +Access. A download would be written there, resumable and reload-proof like any +local file, and handed to the person at the end: `handle.getFile()` returns a +`File` backed by disk, so `URL.createObjectURL` on it is a link the browser +copies disk-to-disk without loading gigabytes into the tab. + +So it is not a browser limitation. It is a **second download target to write +and maintain**, and it costs: + +- a fourth tier in `_openDownloadTarget`'s already four-deep chain — the + function whose silent floor cost this project a gigabyte of RAM per film + once already (`CLAUDE.md`); +- **quota.** OPFS is subject to storage quota and to eviction unless + `navigator.storage.persist()` is granted. A 4 GB film may simply not fit, and + finding that out at 90% is worse than not offering the path; +- **double disk usage and a real wait at the end.** The file exists twice while + the browser copies it out, and the copy is not instant for a large file; +- **cleanup.** Abandoned partial downloads sit in OPFS invisibly, consuming the + origin's quota, until something reaps them. That is a new janitor with its + own failure mode; +- **an API matrix that must be measured, not read.** The OPFS write path is not + uniform: `createSyncAccessHandle()` (workers only) and `createWritable()` have + had different availability across Firefox and Safari versions. This repo's + standing rule applies — *model the environment, never the code under test* — + so the answer comes from a probe in real browsers, not from a compatibility + table. + +**Recommendation.** Not in this work. Revisit it as its own change once transfer +slots are in production, and if it is picked up, it is worth more than pause +alone: it would give Firefox and Safari resumable downloads, reload-proof, and +retire the service-worker path and its iframe entirely. Ship it as a target +first and let pause fall out of it — the reverse order is how a fourth tier +gets bolted onto the chain in a hurry. + +## 7. Settings + +### 7.1 Node-wide — draft v6 §2.11's pattern, exactly + +Two new `[node]` keys, `max_concurrent_downloads` and +`max_concurrent_uploads`, default 8, positive int, `_positive()`-validated like +`max_concurrent_streams` (`config.py:332`). Then, mechanically: + +- `roster.py`: `SETTING_MAX_DOWNLOADS` / `SETTING_MAX_UPLOADS`, added to + `node_settings()` (`roster.py:943`); +- `ops.py`: entries in `get_node_settings` / `set_node_settings` + (`ops.py:1290`, `ops.py:1309`), written to `roster.db` **and** `node.toml`; +- **hot-swap through one helper.** `ops.set_node_settings` currently pokes + `webrtc._stream_sem`, which does not exist (§1). Replace all three with + `webrtc.set_capacity(streams=…, downloads=…, uploads=…)` on + `WebRTCTransport`, which resizes the live pools and pumps the queues. One + function, on the object that owns the state, with a test that changes a cap + on a running transport and watches a queued transfer start; +- the Node page grows two number fields beside "max concurrent streams" + (`node-page.js:1042`) and the CLI grows + `meshbay-node transfers show|set` — plus `meshbay-node transfers list`, + which prints live leases and the queue from `GET /api/transfers`. + +### 7.2 Per group, per member — `apps_enabled`'s pattern, exactly + +`OP_TRANSFER_LIMITS` in `adminop.py`, subject `"d=2,u=2"` so what the operator +signs names the outcome. `_do_transfer_limits` + +`_admin_exec_transfer_limits` in `webrtc_server.py`, stored by +`ops.set_transfer_limits` in `roster.db` group settings, broadcast as +`transfer_limits_ack` to the group's peers, surfaced in the group Settings tab +as a section beside the scan settings. + +**Absent means the default (2), not unlimited.** Deliberately unlike +`member_upload`'s "absent means allowed": a group that predates the setting and +came back unlimited would leave the node-wide cap as the only control, which is +the situation this work exists to end. Enabled by default, breaking, no opt-out +switch — the same call as MNP 2.0's chat encryption, for the same reason. No +release note and no operator prose: the project is in development and every +node is a test node. + +Bounds: 1–32. Zero is not "unlimited" and is not accepted; a member who may not +transfer at all is a member the operator revokes. + +--- + +## 8. The client + +### 8.1 `TransferStore` — three new states, one changed contract + +``` + ┌──────────┐ slot granted ┌──────────┐ + start() ─────▶│ queued │────────────────▶│ running │ + └──────────┘ └──────────┘ + ▲ │ │ │ │ + resume() │ │ cancel() pause() │ │ │ done + │ ▼ ▼ │ ▼ + ┌──────────┐ ┌────────┐ ┌──────┐ + │ paused │◀──────────────│ paused │ │ done │ + └──────────┘ └────────┘ └──────┘ + │ │ + cancel│ failed / cancelled + ▼ + ┌───────────┐ + │ cancelled │ + └───────────┘ +``` + +Statuses become `queued | running | paused | done | cancelled | failed`, and +`run({signal, onProgress})` becomes +`run({signal, onProgress, lease, resumeFrom})`: + +- `lease` — an object the store hands the runner: `await lease.acquire()` + blocks until granted, `lease.release(reason)` in a `finally`. The runner + never speaks to the transport about slots directly. One place opens a lease, + one place closes it — the same rule as the node's. +- **The slot is asked for after there is somewhere to write, and that ordering + is load-bearing.** Asking first reads better — the widget could draw a row + while the target is being chosen — and is wrong: opening a target takes tens + of seconds of streamed-download timeouts, or as long as somebody leaves a Save + As dialog open, and a grant not taken up in time is revoked. Tried the other + way during the build: three downloads started, one arrived. A test pins the + order now. +- The store gains `pause(id)`, `resume(id)`, `pauseAll()`, and a persisted + resume record for the tiers that support it. +- `_maybeRelease(transport)` must treat **paused and queued as busy** + (`transfers.js`'s `_busy`). A paused transfer whose transport was closed + because the group page went away can never resume — that is a one-line + regression waiting to happen, and it has a test. + +### 8.2 The widget + +`TransferWidget` (`app.js:141`) is a flat list of rows with a progress bar. It +becomes a small panel that can answer, at a glance, *what is happening and what +is my fault*: + +``` +┌─ Transfers ───────────────────── 2 running · 3 waiting ─┐ +│ ⏸ Pause all Clear finished │ +├──────────────────────────────────────────────────────────┤ +│ ⬇ Some Saga (2019).mkv ⏸ ✕ │ +│ ███████████████████░░░░░░░░░░ 62% │ +│ 2.4 GB / 3.9 GB · 11.2 MB/s · 2 min left │ +├──────────────────────────────────────────────────────────┤ +│ ⬆ holiday-photos.zip ⏸ ✕ │ +│ ████████░░░░░░░░░░░░░░░░░░░░ 27% │ +│ 118 MB / 430 MB · 3.1 MB/s · 1 min left │ +├──────────────────────────────────────────────────────────┤ +│ ⬇ A Different Show S02E04.mkv ✕ │ +│ ┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈ waiting · 1 ahead │ +│ Your 2 slots are busy │ +├──────────────────────────────────────────────────────────┤ +│ ⬇ render.tar ▶ ✕ │ +│ ██████░░░░░░░░░░░░░░░░░░░░░░ paused at 21% │ +├──────────────────────────────────────────────────────────┤ +│ ⬇ notes.pdf Finished ↗ │ +└──────────────────────────────────────────────────────────┘ +``` + +Concretely: + +- **A summary in the header** — `2 running · 3 waiting` — and the nav badge + counts running *and* waiting, so a queued transfer is never invisible. +- **Waiting rows have a striped, indeterminate bar** (a CSS gradient animation, + no JS), visibly different from a stalled progress bar. `prefers-reduced-motion` + turns the animation off, not the state. +- **The reason for waiting is written out**, from the node's own counters: + *"Your 2 slots are busy"* vs *"The node is at capacity — 1 ahead"*. Someone + should be able to tell their own limit from the operator's without asking. +- **Pause / resume / cancel as icon buttons**, with real `aria-label`s and + `title`s; pause is **absent, not disabled**, where the target cannot resume + (§6.1) — and the row carries a quiet hint saying why, once. +- **ETA** from the existing windowed speed, hidden below 5 s of samples so it + does not flicker nonsense at the start. +- **Ordering**: running, then waiting (queue order), then paused, then + finished. A row does not jump when its neighbour finishes — group headers + rather than a re-sort on every emit. +- **Live region**: state transitions announce once (`aria-live="polite"`), + progress does not. +- Colours from the existing CSS variables, so light and dark both work without + a second palette. New rules go in `style.css` beside `.transfer-*`, measured + with `tests/harness/layout_probe.py` at 320 px — the transfers panel is + *precisely* what that harness was written for, after it hung 138 px off the + left of a phone. + +### 8.3 i18n + +New keys: `transfers.waiting`, `transfers.waiting_ahead`, `transfers.paused`, +`transfers.pause`, `transfers.resume`, `transfers.pause_all`, +`transfers.your_slots`, `transfers.node_busy`, `transfers.eta`, +`transfers.interrupted`, `transfers.no_pause_here`, `transfers.running_count`, +plus the settings labels. `en.js` is the source; `test_locales.py` holds the +other nine to its key set and will fail until all ten are translated. + +--- + +## 9. Regressions to watch + +Ranked by how quietly they would ship. + +1. **Browsing degraded by a busy node — the operator's stated requirement + (§3.4).** `_try_serve_thumbnail` shares `file_req` with real files, and the + Photos viewer and the Files preview modal fetch whole index entries through + `pipelinedDownload` just as a download does. Four tests, not one: with both + pools full and a queue behind them, (a) a Videos poster grid still fills, + (b) a Music album grid still fills, (c) a photo opens full size, (d) a PDF + preview opens. Each must also fail with the exemption removed. +2. **Video streaming charged twice.** `stream_req` must not touch the download + pool. Test: fill the download pool, start a film, assert it plays. +3. **Chat attachments — decided: they take an upload slot** like any other + upload, through the same `_do_file_upload` path. The regression is therefore + in the *interface*, not the rule: a paperclip whose file silently sits in a + queue is a chat that looks broken. `chat-app.js`'s composer must render the + same waiting state as the transfers panel — the attachment row says + "waiting, 1 ahead", not nothing — and a queued attachment must not block the + composer, which is the freeze already recorded in `CLAUDE.md` for chat + sends. Test it in `chat_send_probe.py`, which drives composer and transport + together for exactly this class of seam. +4. **`_maybeRelease` closing a transport with paused transfers on it** (§8.1). +5. **Zip downloads.** One lease for the whole job. Getting this wrong means a + 40-file zip takes 40 slots and deadlocks against its own cap — a deadlock, + not a slowdown, because the job cannot finish until it holds them all. +6. **Search across groups** (`search-page.js`, `getTransport`) opens transports + to several nodes; each node caps independently. Correct, but the widget must + not present another node's queue as this one's. +7. **`_free_name` and resumed uploads.** A resumed upload must reuse + `state["stored_name"]`, not draw a new one — otherwise a reconnect produces + `file (2).mkv` next to a half-written `file.mkv.part`. +8. **The 60 s cap in `_do_file_request`'s backpressure loop** interacts with a + paused reader: a paused transfer stops requesting, so nothing is waiting — + but check that a *slow* reader is not now reclaimed as idle. The idle timer + is reset on each `file_req` under the lease, not on each chunk *sent*. +9. **`test_security_regressions.py:786`** asserts `_transcode_sem` is in the + source. Renaming it during the `set_capacity` refactor will fail that test — + which is the test doing its job. Update it deliberately. + +10. **The one this list did not anticipate, and the one that happened:** none of + the eight defects the live pass found are in this list, because every entry + here is about *slots* and seven of the eight were about the download path + underneath them (§12.2). A regression list written from the change being + made will not cover the ground the change stands on. Before the flag day, + walk the download path itself — every write target, on every browser — + rather than the diff. + +--- + +## 10. Tests + +Following the repo's rule: measure the environment, never model the code under +test. + +**Node (`packages/meshbay-node/tests/`)** + +- `test_transfer_slots.py` — real `WebRTCPeerSession` objects against a fake + DataChannel (the shape `test_webrtc_transport.py` already uses): node cap, + per-member cap across two sessions of one account, queue order, head-of-line + skip, grant deadline, idle reclaim, auto-close on the last chunk, release on + `shutdown_tasks`, no double-charge on a re-`transfer_open` with the same `tr`. +- `test_transfer_invariant.py` — the randomised property test of §5.3. +- `test_transfer_exemptions.py` — the §3.4 requirement, as behaviour: posters, + covers, thumbnails, a full-size photo and a document preview all served with + both pools full and a queue waiting; the 2-in-flight bound refusing a third + concurrent leaseless entry; and a source-reading check that the only callers + opening a lease are the three that call `transfers.start()`. +- `test_navigation_not_leased.py` — `index_sync`, `index_delta`, + `media_meta_req`, `music_meta_req` and directory listing answered normally + with both pools full. Cheap, and it fails the day someone routes a listing + through `file_req`. +- `test_upload_resume.py` — reconnect mid-upload resumes at `next_index`; + orphan `.part` reaped; `stored_name` preserved. +- `test_node_settings.py` extension — `set_capacity` actually resizes a live + pool (the test `ops.py:1362` never had). + +**Hub / SPA (`packages/meshbay-hub/tests/`)** + +- `test_transfers.py` extension, under Node as today: queued → running, pause + releases, resume re-queues, cancel while queued, `_busy` counts paused. +- `tests/harness/transfers_probe.py` — mounts the real `TransferWidget` in + Chrome against a stubbed store and reads the rendered states back. The + precedent is `chat_send_probe.py`: the seam between store and widget is where + this will break, and neither source shows it. +- `test_layout_measured.py` extension — the panel at 320 px with a waiting row. +- `test_locales.py` — passes only when all ten catalogues have the new keys. + +**Live — `packages/meshbay-node/tests/transfer_probe.py`, in the repo** + +Not in `QE/` as this section first said: it found four defects nothing else +could, and `QE/` is not versioned, so it lived on one machine. It is not +collected by pytest (the filename does not match `test_*.py`) and still needs +`QE/deploy/e2e.py` and `demo.env` at run time, which it locates and explains +rather than importing blindly. + +`--want N` measures the cap, `--pull N [--parallel]` downloads real files to +completion on one connection, and `--operator` covers the two things only the +operator's CLI can answer: that a cap raised live starts what was waiting, and +that a vanished peer's slots are back before anyone asks. + +- ~~`QE/deploy/transfer_probe.py`~~ — opens N real MNP sessions as two accounts, + starts more transfers than the caps allow, and asserts the observed + concurrency and the order they complete in. The equivalent of + `stream_probe.py`, and the only thing that answers "is it the node or the + browser" in one run. +- A two-browser manual pass: cap 2, start 5, watch the queue drain; pause one, + reload the page, resume it; kill a tab mid-transfer and watch the slot come + back within 120 s. +- `QE/migration/reap_orphan_parts.py` — one-shot, node stopped, for the `.part` + files already on deployed nodes (§6.4). + +Per the QE rule in `CLAUDE.md`: any test against meshbay.org opens the UFW +port, tests, and closes the port and kills the processes **in the same block**. + +--- + +## 11. Order of work + +Each stage is shippable and testable on its own; nothing before stage 5 changes +what a member sees. + +| # | Stage | Contents | +|---|---|---| +| 1 | ✅ **Capacity, fixed** | `set_capacity()` on `WebRTCTransport`, replacing the dead `_stream_sem` poke. No new features. Ships the §1 bug fix alone, where it can be verified alone | +| 2 | ✅ **Leases, node-side** | Lease table, pools, queue, pump, every reclaim path in §5.1, `GET /api/transfers`, the DEBUG line. No client uses it yet; the node grants everything because no client asks | +| 3 | ✅ **Settings** | `[node]` keys, roster overrides, `OP_TRANSFER_LIMITS`, Node page fields, CLI verbs, ack field | +| 4 | ✅ **MNP 3.0** | Version bump on both sides, `MNP_MIN_SUPPORTED` at 3.0, **§3.4.1's bound on leaseless reads** (which did not exist), the desktop client's version gate, every package aligned on 0.13.0 | +| 5 | ✅ **Client leases** | `TransferStore` states and the lease contract, `file_req`/`file_upload` carrying `tr`, reconnect re-open, watchdog | +| 6 | ✅ **Widget** | The panel of §8.2, i18n, layout measured | +| 7a | ✅ **Pause / resume, in session** | Per-target `pausable`, the slot released on pause and re-asked on resume, `fromChunk` in the pipeline, the widget's pause button | +| 7b | ⬜ **Resume across a reload** | The persisted resume record, `save:resume` in Electron, `createWritable({keepExistingData:true})` with a position, truncate-to-chunk | +| 8a | ✅ **Upload state + reaping** | `uploads.py`, state in the group context keyed by member, the `.part` janitor | +| 8b | ✅ **Upload resume + pause** | The sealed probe chunk, `resume_from`, uploads asking for their own slot, `touch()` on an upload chunk | +| 9 | ✅ **Live pass** — done out of order, and §12 is its report. | `transfer_probe.py`, two browsers, the manual list | +| — | *Deferred, separate change* | **OPFS as a download target** for Firefox and Safari (§6.5) — resumable downloads there, and the retirement of the service-worker path. Not part of this work | + +Stages 1–3, 5 and 6 have landed. Stage 9 was run **before** stage 4 rather than +after, deliberately: while nothing is enforced, a defect in the machinery costs +nothing, and §12 is what that bought. Stage 4 is a flag day and should land only +on the evidence §12.3 describes. + +--- + +## 12. What the live pass found, and what it changes + +Step 9 was meant to confirm the machinery. It found **eight defects**, seven of +which no test in the repo could reach, and three of which are in code the +sections above describe as settled. Four more (§12.4 to §12.7) came out of +fixing those eight, and the last of them — a hard reload silently disabling the +only unbounded write path Firefox has — was the one actually being reported all +along. They are listed here because their pattern matters more than +any one of them: every one needed a real browser, a real node and a real hub to +show itself, and the last two were only reachable on Chrome. + +### 12.1 The three in the lease machinery + +| | Found by | +|---|---| +| **`touch()` was never called.** The node ignored `tr` on `file_req`, so `used` stayed False for every download and each grant was revoked at 30 s mid-transfer | the node's own log | +| **The requeue was a permanent cycle.** Revoked → queued → granted again → revoked, every 30 s, for as long as the daemon ran | the node's own log | +| **`transfers show` reported the module defaults**, so `transfers set 2 2` answered "applied now" and the next line said 0/8 | typing the command | + +The first is the one to learn from. `TransferSlots.touch()` existed, was +covered by its own test, and **had no caller**. The pool was right, the +handlers were right, and the call between them was missing — so neither side's +tests could see it, and the property test could not either: nothing drifted. +**A seam is not tested by testing both sides of it.** + +### 12.2 The five in the download path itself + +None of these are about slots, and all of them would have become "this download +is impossible" the moment leases were compulsory: + +- **three headers** decide whether the page may frame its own `/_mbdl/<id>` URL + — `frame-src`, `frame-ancestors`, `X-Frame-Options` — and all three were + wrong. Fixed one at a time over an afternoon; all three were visible in a + single `curl -I` against the deployed hub; +- **a service worker with no event for ~30 s is terminated**, and a streaming + response does not count as activity. The reader vanished mid-file and + `writable.write()` never resolved *and never rejected*: no error, no log, a + progress bar that stopped, and a node that stayed healthy throughout; +- **`encodeURIComponent` leaves `'` alone** and `'` is RFC 5987's delimiter, so + a 449 MB film arrived complete and correct under the name + `mtsshk9w-ohqty535`; +- **a browser grants one file picker per user gesture**, so downloading three + files at once failed two of them with a message about gestures; +- **two silent returns** meant a click with no connection produced nothing at + all — no transfer, no icon, no message. + +### 12.3 What this changes about step 4 + +Step 4 makes leases compulsory and refuses 2.x peers at the handshake. It is the +only irreversible step in this plan. + +Every defect in §12.1 and §12.2 was invisible to 1169 node and 814 hub tests, +and was found by a person clicking Download and pasting a console. Three of them +were introduced *by this work* on the day it was written. Making a mechanism +compulsory is a bet that it is well understood, and the evidence of one +afternoon is that it was not. + +**The recommendation was to wait for evidence rather than for time**: a week of +ordinary use with no freeze and no lost transfer, or the equivalent in deliberate +runs of `transfer_probe.py --pull 3 --parallel` and `--operator`. + +**Overruled by the operator on 2026-09-09, and the reasoning is better than the +recommendation it replaced.** The concern was that a flag day removes the +fallback exactly when defects are still being found; the answer is that the +fallback only has value while more work is coming down the same path, and none +is — the next work is a music application with playlists, which touches none of +this. The prerequisite this section named was satisfied first: +`GET /v1/hub/version` already carried `client.minimum`, and the desktop client +now reads it before connecting instead of meeting a handshake refusal it has no +vocabulary for. + +Nothing about §4's reasoning has changed. What changed is the confidence that +the thing being made compulsory works. + +### 12.4 The silent row — fixed + +**The transfers panel showed nothing while the target was being opened.** The +row was created after `_openDownloadTarget` returned, which can take tens of +seconds, so three clicks produced no icon at all and then several rows at once. + +It was "fixed" during the build by taking the lease first, which caused §12.1's +`not_taken_up` symptom and was reverted — on a wrong diagnosis, as it turns out: +the revocations were the missing `touch()`. The revert was right anyway (§8.1). + +The shape that worked is a `prepare` step in the store, distinct from `run`: the +row appears at the click, the target is opened, and the lease is asked for last. +The reservation about `showSaveFilePicker` needing a user gesture turned out to +be real, but not in the way expected — see §12.5. + +### 12.5 One dialog per file, and three downloads frozen behind it + +Reported from Chrome, after §12.4 shipped: selecting four files produced a Save +As dialog for the first, then — once that file had finished — a dialog for the +second, while the last two timed out. On a later attempt the three remaining +transfers simply froze. + +Three facts explain it, and only the third was a surprise: + +1. `for (const entry of selected) await downloadFile(entry)` used to serialise + the target openings **by accident**. Moving the opening into `prepare` + removed the accident and four pickers raced. Fixed by `_openTargetInTurn`, + an explicit queue on the targets — never on the rows, so every download + still appears at the click. +2. A browser grants one file picker per user gesture, and selecting four files + is one gesture. The code already recovered from the `SecurityError` Chrome + throws for a picker with no gesture behind it, by streaming instead. +3. **Chrome does not throw.** It shows the dialog anyway and waits for a human. + So the recovery in (2) was never reached, and the queue from (1) turned the + unanswered dialog into a head-of-line block: the third and fourth downloads + were not frozen, they were waiting correctly for a dialog nobody had been + shown yet. From the panel that is indistinguishable from a freeze. + +The fix is to stop asking. `_openTargetInTurn` marks everything that has to wait +its turn as `batched`, and a batched opening prefers the streamed path whatever +the download mode says. The first file of a batch — the one that actually holds +the gesture — still gets its dialog, so the preference is honoured where it can +be. For the rest there is no gesture left to spend, so nothing is lost by +streaming: the file still lands on disk, in the browser's own download folder, +written as it arrives. Only the choice of folder goes, and it was not on offer. + +If the worker does not answer, a batched download falls back to the dialog +rather than failing — asking is better than losing the file (§ "a preference +must not cost a capability", which applies to the fix as much as to the bug). + +Pinned by `test_targets_are_opened_one_at_a_time` (the queue, and that the first +opening is the only unbatched one) and by four cases in `test_memory_ceiling.py` +(the branch itself, both fallbacks, and that batching never pushes a large file +into memory). + +**What that left, and what was measured.** With the queue unblocked all four +files downloaded, but a dialog still appeared for each one — with MeshBay's own +setting on "save automatically", which should never prompt. That points at the +streamed path failing, so it was measured rather than reasoned about: a real +Chrome 152 driven over CDP against the deployed hub, running the actual flow +(a `TransformStream` posted to the worker, a hidden iframe on `/_mbdl/<id>`, a +wait for `mbdl-serving`). It is served in 2–3 ms on a normal load, after a hard +reload, and twice in the same document. Ctrl+F5 does leave the page +uncontrolled — `navigator.serviceWorker.controller` is null and no +`controllerchange` arrives — but the `mbdl-claim` recovery already in +`_claimController` gets control back inside 3 s. Both probes are in +`QE/`-style scratch scripts, not in the suite: they need the deployed hub. + +**And the queue itself became the next defect.** Serialising the openings was +new in the same commit, and on Firefox it regressed what had always worked: +four downloads that opened their targets at the same time began waiting on the +slowest, and all four sat at "preparing" — the node journal showing +`d=0/8(q0) u=0/8(q0)`, not one transfer opened, so the block was entirely in +the client before any slot was asked for. Measured on Firefox 154 against the +deployed hub: `register` and `ready` return instantly, the page is controlled, +and four serialised openings are served in 5–18 ms — so the streamed path was +not the delay either. The queue was. + +Two bounds fix it, and both are narrowings of the queue rather than of any +capability. Only an opening that could actually show a dialog joins the queue, +which on Firefox and Safari — no `showSaveFilePicker` at all — means none of +them do, restoring exactly the previous behaviour. And no opening waits behind +another for longer than `TARGET_QUEUE_BUDGET_MS` (90 s), because `_targetQueue` +is never reset and an opening that never settles would otherwise leave the page +unable to start a download again until it is reloaded. Releasing early is safe: +whatever is ahead is still the only unbatched opening, so the released one takes +the streamed path and opens no second dialog. + +The general shape, third time in this section: **a queue is a way to convert one +slow participant into several stuck ones.** Every queue added here needs to say +what it protects and refuse everything it does not protect, and needs a bound. + +### 12.6 The wait with no deadline + +Bounding the queue was not enough: Firefox still showed four rows at +"preparing", with the queue bypassed, so each opening was hanging on its own. + +`_claimController` had two waits with no deadline of any kind — +`navigator.serviceWorker.register()` and `navigator.serviceWorker.ready` — +while `SW_CONTROL_BUDGET_MS` bounded only the wait that comes *after* them. And +`_swPromise` is shared, so a single unsettled one of those left every download +on the page waiting on the same promise, for the life of the page. + +Measured on Firefox 154, on a local `127.0.0.1` site so no hub was involved: + +| worker | `register()` | `ready` | +| --- | --- | --- | +| that installs | 8 ms | 0 ms | +| **whose install handler rejects** | **7 ms** | **never settles** (still pending at 10 s) | + +That is the whole mechanism. `register()` resolves as soon as the registration +object exists — carrying nothing but an *installing* worker — and `ready` is +what waits for an active one. A worker that cannot install therefore produces a +registration that looks fine and a `ready` that never comes. + +Every wait in `_claimController` is now inside one budget, with two carve-outs +that exist so a deadline never costs a capability. A `ready` that times out +while `registration.active` is set is not fatal — `ready` may be waiting on a +newer worker that cannot install while an older one serves perfectly well. And +the `mbdl-claim` recovery keeps its own budget outside the deadline, because +giving up there would cost Firefox the only unbounded way it has to write a +download to disk. + +**A deadline alone would still not have been a fix**, only a better-explained +failure: a registration stuck with nothing but an installing worker does not +heal, and every later visit finds the same one. So when `ready` times out with +no active worker, the registration is discarded (`unregister()`) and asked for +once more with a fresh budget. The page repairs itself instead of needing +developer tools. + +Pinned by four cases in `test_streamed_download_reliability.py`, each checked +against the unfixed source: a worker that never installs, a registration that +never answers, a stuck `ready` that must not throw away a working worker, and +the discard-and-retry. + +So the streamed path is not what is prompting. Exactly three things can open a +dialog per file on Chrome in automatic mode: our `showSaveFilePicker`, the +folder permission bubble from `ensurePermission` (only when a folder was granted +and the permission lapsed, and it is not a Save As), and **Chrome's own +"always ask where to save each file"** in `chrome://settings/downloads` — which +applies to the worker's response because it carries `Content-Disposition: +attachment`, and which no download mode of ours can override. The `console.info` +added with this fix distinguishes them in one line: it is written before every +dialog we open, so a dialog with nothing in the console is the browser's. + +### 12.7 The one none of the above explained: a hard reload + +Every fix in §12.5 and §12.6 was real, and none of them was the defect being +reported. Downloads on Firefox failed with "the worker did not answer the +download within 15s", every time, for the operator, while the *same profile* +driven from this machine — headless and headed, on their display, with their +files, through the real interface — succeeded every time. + +The operator's own sequence found it, and it is worth writing down verbatim +because no automated run could have produced it: a freshly started browser +downloaded four files out of four, twice; one **Ctrl+F5** and every attempt +afterwards failed; restart, fine again; Ctrl+F5 before any attempt and the very +first one failed. + +**A document fetched by a hard reload is loaded with the service worker +bypassed.** It can still be claimed afterwards — so +`navigator.serviceWorker.controller` comes back and every check in +`_claimController` passes — but the navigations that document starts keep +missing the worker, and the hidden iframe a streamed download needs *is* a +navigation. On Firefox and Safari that is the only way to write a file too +large to hold in memory, so the download cannot happen at all, for the life of +that page. The same behaviour had already been measured on Chrome in §12.6's +work and its significance was missed. + +The hard reloads were on this author's instruction, after each deployment. The +SPA's HTML is served `no-store`, so an ordinary reload has always picked up a +new build and Ctrl+F5 was never needed for anything. + +**Why the measurements could not find it.** WebDriver cannot perform a hard +reload — the key event goes to the content, not the browser chrome — so every +automated run tested, over and over, the one case that works. A fix is validated +by tests; a diagnosis is not always validated by automation, and the operator's +manual sequence should have been asked for hours earlier. + +The remedy is to stop inferring servability from control and to ask instead. At +boot the client opens a four-byte stream and a hidden iframe, exactly as a real +download would, and tears both down. If the worker does not answer, the page +reloads **once**, ordinarily, which puts it back under the worker; the flag is +in `sessionStorage` because it has to survive the reload it triggers and must +stop rather than loop. Two delays then had to go, because the repair first +landed about thirty seconds late — long enough to click and watch four rows +hang: the claim is asked for *before* the control budget rather than after it +(a page uncontrolled beside an active worker is never claimed on its own), and a +download that starts while the self-test is running waits for it instead of +racing it. + +And the refusal message no longer tells a Firefox user to switch to Chrome for a +state an ordinary reload undoes. All ten catalogues say to reload first. + +--- + +## 13. Decisions — all settled (operator, 2026-09-08) + +Nothing in this plan is waiting on an answer. + +| # | Decision | Consequence | +|---|---|---| +| 0 | **Browsing is never subject to a slot** — posters, thumbnails, covers, previews, listings | §3.4, structural: a transfer is what the widget shows | +| 1 | **MNP 3.0, hard refusal of 2.x at the handshake** | No leaseless branch survives anywhere. Flag day, stage 4 | +| 2 | **Per-member default of 2 applies to existing groups** | No release note, no migration prose — the project is in development and every node is a test node | +| 3 | **Chat attachments take an upload slot like any other upload** | §9.3: the composer must show the waiting state, or the paperclip looks broken | +| 4 | **Pause on the service-worker path (Firefox, Safari) is deferred, not abandoned** | §6.5 says what it would actually take. It is cost, not impossibility | +| 5 | **The two big caps stay node-wide** | The resource is the machine's. A per-group variant can be added later with no protocol change | + |