# Transfer slots, queueing, pause and resume — build record > **Status: built.** Leases, caps, queueing, the settings surface, the client > widget, pause and resume all shipped; the compulsory-lease flag day is > **MNP 3.0**. > > **The design is `MESHBAY_DESIGN.md` §5.5**, which states the lease model, the > caps, and the reasoning behind the browsing exemption. Read that first. This > document is kept for what a synthesis cannot carry: the failure-mode analysis > in §5 (every way a slot can be lost, every way a client can be left waiting, > and the invariant that checks both), the target-by-target resumability table in > §6, and §12 — what a live pass found after the work was called done, which is > the most useful part of the file. > > **Where this document and the code disagree, the code is authoritative**, and > two places are known to disagree: the leaseless-read bound is 12 files per > session with an idle expiry, not the 2 written in §3.4.1, and the wire section > describes MNP 3.0 as pending when it has shipped. `MESHBAY_DESIGN.md` §5.5 > carries the current values. > > Convention carried over and not negotiable: **a claim here names the failure it > holds against.** "A slot is released when the transfer finishes" is not a > design; "here is every way a slot can be lost and what reclaims it in each > case" is. --- ## 0. The short answer | Asked for | Feasible | Where the work is | |---|---|---| | `max_concurrent_downloads` / `max_concurrent_uploads` per node (default 8) | Yes | The node has **no concept of a transfer** today — that is the whole job | | Per-member caps inside a group (default 2) | Yes | New per-group operator-signed setting, same shape as `apps_enabled` | | "Max reached" refusal + client queues the rest | Yes | Better as a **node-side queue with an explicit grant**, not a client retry loop | | A visible **waiting** state in the transfers widget | Yes | `TransferStore` already has a status field; it gains three states | | Cancel | **Already works** (`transfers.cancel`, `transfers.js`) | Needs to also tell the node, which today it does not | | Browsing — posters, thumbnails, previews, listings — untouched by the caps | Yes | §3.4. Structural, not a threshold: the three functions that draw a widget row are the three that take a slot | | Pause with resume | Yes, **in three tiers** — see §6. Not every download target can be resumed, and the interface must not offer what a target cannot do | The risk is not in any one of these. It is that a capped resource plus a queue is exactly the shape of bug that shows up as *"my download says waiting and never starts"*, days later, on someone else's machine. §5 is therefore the longest section in this document and is the one to review hardest. --- ## 1. What exists today, measured Read this before designing anything: three of the assumptions one would naturally make are false here. **A download is invisible to the node.** `pipelinedDownload` (`static/file-utils.js`) sends `PIPELINE_WINDOW = 8` independent `file_req` messages and reassembles the answers. Each one is served by `_do_file_request` (`webrtc_server.py:3380`), which looks the entry up, reads and encrypts one 1 MB chunk, waits for room on the channel (`DOWNLOAD_BUFFER_HIGH`, `webrtc_server.py:253`) and sends it. **The node never learns that a download started, and never learns that one ended.** There is nothing to count and nothing to cap. This is the central fact of this work. **An upload is half-visible.** `_do_file_upload` (`webrtc_server.py:4471`) keeps `self._uploads[f"{rel_dir}/{filename}"]` with `next_index` and a `.part` file on disk. That state is **per session object and in memory**: a browser that disconnects mid-upload leaves a `.part` file on the operator's disk that nothing ever removes, and a retry starts at chunk 0 under a new `_free_name`. So there is already a resume story here, it is just not reachable — and there is already a leak. **Video streaming is capped, and its cap is a good model.** `_stream_video` (`webrtc_server.py:5230`) takes `ctx["_transcode_sem"]`, refuses with `"Server busy, retry shortly"` when it is empty, and every hard-won lesson in this repo about slots — `_replace_stream`, `shutdown_tasks()`, `_spawn()` and the garbage-collected task, `await proc.wait()` after `kill()` — is a lesson about *not losing a slot*. That history is why §5 exists. **One bug found while reading, in scope for this work.** `ops.py:1362-1365` hot-swaps the stream cap by assigning `webrtc._stream_sem` — an attribute that does not exist. The real semaphore is `ctx["_transcode_sem"]` (`webrtc_server.py:5220`), and `hasattr(webrtc, '_stream_sem')` is always False, so **changing `max_concurrent_streams` from the Node page has never taken effect without a restart**, contrary to §2.11 of draft v6. Two more hot-swappable caps are about to be added next to it; fix it with one shared helper rather than three copies of the same mistake. **Existing client state.** `TransferStore` (`static/transfers.js`) already survives leaving a group, already holds the transport open until the last transfer finishes (`releaseWhenIdle`), already cancels on sign-out, and already computes a windowed speed. Its statuses are `running | done | cancelled | failed` and its `run()` contract is a promise plus a polled `signal.aborted`. That contract is the thing that has to change, and it is the only client-side structural change. --- ## 2. What a naive implementation gets wrong Worth stating, because each of these is a plausible first attempt. 1. **Counting `file_req` messages.** A rate limit on chunk requests caps throughput, not concurrency, and gives the client nothing to render as "waiting". It also cannot distinguish a download from a thumbnail. 2. **Inferring a transfer from activity.** "A (session, file) pair with a chunk request in the last 30 s is a live download" needs no protocol change and is tempting. It makes the admission decision unobservable — the client cannot be told *why* it is being refused, cannot be told when to try again, and a paused transfer is indistinguishable from a crashed one. 3. **A client-side queue only.** The client already knows what it wants to download; it does not know what the other five members are doing. A queue held only by clients cannot order anything fairly and turns into a retry storm against a busy node. 4. **A refusal the client retries on a timer.** This is the version that produces the "stuck at waiting" report: the client polls, the node refuses, nobody is at fault, and there is no evidence anywhere. The node must **push** the grant. 5. **Gating every `file_req`.** Thumbnails, posters, cover art and audio transcodes are served through the *same* `file_req` path (`_try_serve_thumbnail`, `webrtc_server.py:3353`), and so are image and text previews. Requiring a slot for all of them means opening a poster grid queues behind a film. §4.4 handles this. 6. **Holding a slot while paused.** If pausing keeps the slot, one member can pause eight downloads and close the node to everyone else. Pausing must release. --- ## 3. The design ### 3.1 The lease Introduce **the transfer lease**: the node's record that a peer is transferring something, held for the length of the transfer and released by name. ``` transfer_open { tr, kind: "download"|"upload", bytes, chunks, label_hint } → transfer_state { tr, state: "granted" } → transfer_state { tr, state: "queued", ahead: 3 } …later… → transfer_state { tr, state: "granted" } ← pushed, not polled transfer_close { tr, reason: "done"|"cancelled"|"paused"|"failed" } → transfer_state { tr, state: "closed" } ``` Six properties, each of which is load-bearing: - **`tr` is drawn by the client**, 16 random bytes hex, exactly like `upload_id` (`transport.js:2208`). Re-opening after a reconnect with the same `tr` is idempotent, so a reconnect cannot double-charge a member for one transfer. - **A lease is scoped to the connection, never to the account.** It dies with the session, which is what makes the primary reclaim deterministic (§5). - **A lease covers a job, not a file.** A directory zip (`downloadDirectory`, `file-utils.js`) is dozens of files and is **one** lease. So is a resumed transfer. - **A lease is never persisted.** Not in `roster.db`, not on disk, not across a node restart. A restart drops every session anyway; a lease that outlived the process would be a slot nothing can release. - **The node counts leases, not bytes.** What a slot protects is concurrency — open file handles, disk seeks, and the channel buffer each transfer keeps full — not bandwidth. - **Every `file_req` for a real file carries its `tr`.** The exceptions are in §4.4 and they are the delicate part of this whole design. ### 3.2 Two caps, in one order | Cap | Scope | Default | Where it lives | |---|---|---|---| | `max_concurrent_downloads` | node-wide, all groups | 8 | `[node]` in `node.toml`, override in `roster.db` (§7) | | `max_concurrent_uploads` | node-wide, all groups | 8 | idem | | `max_downloads_per_member` | one group, one account, **across all their devices** | 2 | per-group, operator-signed (§7) | | `max_uploads_per_member` | one group, one account, across devices | 2 | idem | Downloads and uploads have **separate pools**, and video streaming keeps its own third pool (`max_concurrent_streams`) untouched. A member watching a film is not charged a download slot, and a download does not make the next film answer "server busy" — those are different resources with different costs and merging them would make both caps meaningless. **Order of checks, and it matters:** per-member first, then node-wide. A member at their own cap queues *behind their own transfers* and never consumes a node-wide slot they would then hold while a second member has none. Reversed, one member arriving first takes all eight. **"Per member" means per account, summed across their devices**, resolved with `_sessions_of(user_id)` (`webrtc_server.py:3325`) — which exists for exactly this reason, since device linking landed. Two browsers and a desktop client signed in as the same person share the two slots. Anything else makes the cap a function of how many tabs someone opens. ### 3.3 The queue One FIFO per pool, per node, holding `(tr, session, user_id, group_id, enqueued_at)`. - **Drained in one place.** `_release_slot()` is the only function that returns a slot, and it is the only caller of `_pump_queue()`. Every path that ends a transfer goes through it, in a `finally`. Two functions that both release would be this repo's flow-control-accounting lesson (`CLAUDE.md`) one feature later. - **Head-of-line blocking is skipped, not waited on.** When a node-wide slot frees, walk the queue and grant to the first entry whose *member* is under their own cap. Granting strictly in order would let one member at their personal cap stall the whole node. - **A grant has an acceptance deadline.** 30 s. If no `file_req` bearing that `tr` arrives, the grant is revoked (`state: "queued"` again, at the tail) and passed on. Without this, a client that dies between the grant and the first chunk holds a slot until the idle timeout. - **The queue is bounded**: 32 entries per account per pool. Beyond that the node answers `too_many_queued` and the client holds the rest in its own list, sending `transfer_open` as its own transfers finish. Unbounded queues are how a node runs out of memory politely. - **Positions are pushed, throttled.** After each pump, recompute positions and send `transfer_state {state:"queued", ahead:n}` only to the sessions whose number changed, at most once every 2 s per session. ### 3.4 What is *not* gated — posters, thumbnails, previews, navigation **Operator decision, 2026-09-08: browsing a group is never subject to a transfer slot.** Not the poster grid, not the album covers, not the video thumbnails, not the file list, not opening a photo or a PDF to look at it. A member must be able to browse a group that is at capacity exactly as they browse an idle one. This is a requirement, not a tuning parameter, and the sections below are written to satisfy it structurally rather than by choosing a lucky threshold. **Navigation proper never touches this path at all.** The file list, the directory tree, the poster metadata and the album metadata travel as `index_sync`, `index_delta`, `media_meta_req`, `music_meta_req` and `link_preview_req` — different message types, sealed under the group key, with no relationship to `file_req`. Nothing in this design can reach them. That half needs no rule; it needs only to be verified by a test that fails if someone later routes a listing through `file_req`. The `file_req` path carries three genuinely different things, and they are distinguishable **structurally**, by what the id resolves to and by which function asked: | What | Call site | Resolves to | Rule | |---|---|---|---| | Thumbnails, TMDB posters, cover art, cached audio transcodes | `MediaThumb` (`video-app.js:204`), and the Music/Photos grids through it | a **media-cache id**, not an index entry — `_try_serve_thumbnail`, `webrtc_server.py:3353` | **Never leased, never counted, never queued.** One chunk each, out of a bounded cache the node built itself | | Looking at one file — a photo opened full size, a PDF, an image, a text file | `PhotoViewer` (`photos-app.js:171`), the Files preview modal (`files-app.js:610`) | a real index entry, fetched whole | **Not leased.** Bounded by §3.4.1 below, which no real viewer ever reaches | | Downloading, and uploading | `downloadEntry`, `downloadDirectory` (`file-utils.js`), `uploadFile` (`transport.js:2192`) | a real index entry | **Leased.** These are exactly the three call sites that go through `transfers.start()` — the three that produce a row in the transfers widget | The last column is the whole rule, and it is worth stating as a sentence someone can check by reading: **a transfer is something the transfers widget shows. If it does not appear in that panel, it does not take a slot.** The two sets are the same three functions, which is what makes this verifiable rather than a matter of judgement at each new call site. #### 3.4.1 The bound on leaseless reads A `file_req` with no `tr` on a real index entry is served, subject to one limit: **at most 2 distinct file ids in flight leaselessly per session**, with no queue — the third is refused with `transfer_required`. Why this shape, and not the byte budget an earlier draft of this document proposed: - **A viewer looks at one file.** The photo viewer shows one photo, the preview modal one document. Two is already one more than any of them needs, and is there so that prefetching the next photo stays possible. - **A size threshold does not work here.** A RAW photo out of a camera is 60–80 MB and is *browsing*; a 40 MB archive is a *download*. Size does not separate them, and any threshold that let the photo through would let the archive through too. What separates them is which function asked. - **A byte-rate budget does not work either.** It would have to be large enough for that same RAW photo, at which point it is large enough to be a download channel. Concurrency is the thing being rationed, so concurrency is what the exemption is expressed in. - **What it costs.** A client that lies — labelling a bulk download as a view — gets two files at a time per session instead of its member cap. That is the residual, it is bounded, it is audited (`file_download` already goes to `audit.db` on chunk 0), and it is the same class of statement as the cap itself: **this is a fairness control among cooperating clients**, in the company of `max_concurrent_streams`. It is not a defence against a member determined to saturate the node's disk, and must never be described as one — that member is a member, and the answer to them is `member revoke`. #### 3.4.2 The consequence for the interface A preview never shows "waiting", because a preview never queues. If the 2-in-flight bound is somehow reached, the modal reports a plain error and the person tries again — it does not silently become a queued transfer in a panel they were not looking at. ### 3.5 Where the numbers are visible The node answers `transfer_state` with the current picture — `used`, `cap`, `ahead` — so the client can say *"waiting — 2 of 2 of your slots are busy"* rather than a bare spinner. The same counters go to: - the loopback API, `GET /api/transfers`, so the operator can see live leases and queue depth from the CLI and the Node page, and - a DEBUG line every 30 s: `transfer: d=3/8 u=1/8 queued=2 (skipped=1)`. That line is the `client_diag` lesson applied here: when someone reports a transfer stuck at "waiting", this is the only thing that will say whether the node ever had them in a queue. --- ## 4. Wire protocol — MNP 3.0 ### 4.1 New message types (`meshbay_common/protocol.py`) | Type | Direction | Carries | |---|---|---| | `transfer_open` | client → node | `tr`, `kind`, `bytes`, `chunks`, `from_chunk` | | `transfer_close` | client → node | `tr`, `reason` | | `transfer_state` | node → client | `tr`, `state`, `ahead`, `used`, `cap`, `scope` | `transfer_state` is the only reply, for granted, queued, revoked and closed alike. One message type with a state field, rather than four types, because a client that must switch on the type to find out it is still waiting is a client that will get one branch wrong. `file_req` gains an optional `tr`. `file_upload` gains `tr` beside the `upload_id` that is already there in clear. **Sealing.** `transfer_open` / `transfer_close` / `transfer_state` carry no file names and no paths — `tr` is opaque, `bytes` and `chunks` are numbers — so they stay in clear like `index_progress`, and for the same stated reason. `label_hint` is **not** in the protocol: the client already knows what it named the transfer, and putting a filename on the wire in clear to make a log line prettier is exactly the trade `groupbox.py` exists to refuse. ### 4.2 Why this is MNP 3.0 and not an additive 2.x The messages are additive; **the requirement is not**. A 2.0 client downloading a 4 GB film sends no `tr`, is treated as a leaseless read, and is refused as soon as it opens a third file (§3.4.1) — or, worse, is *not* refused and transfers outside every cap. An opt-in switch — "enforce leases only for clients that speak 3.0" — leaves a leaseless branch reachable on every node, which is finding C6's lesson (`TCP accepted a bare JWT`) one feature later, and it was already refused once for chat encryption on 2026-09-07. So: **`MNP_VERSION = "3.0"`, `MNP_MIN_SUPPORTED = "3.0"`**, and a 2.x peer is refused at the handshake with `version_too_old` rather than admitted and then mysteriously unable to download. Same play, same reasoning, and `handshake.py`'s version range is the mechanism that already exists to pay for it. **What that costs, stated plainly.** The SPA deploys with the hub, so browsers get the new client. The **desktop client ships its own UI**, so an un-updated app is locked out until its user updates — the version-skew consequence draft v6 flagged under "Shipping the UI in a package". Before this ships, `GET /v1/hub/version` needs its minimum-client-version field populated and the client needs to show a real message ("this version can no longer connect, update here") rather than a handshake refusal. ### 4.3 Handshake ack Two fields inside the sealed configuration block (`_complete_handshake`, `webrtc_server.py:852`), so the interface can draw correct numbers before anything is transferred: ```python "transfer_limits": {"downloads": 2, "uploads": 2}, # this member, this group ``` Absent reads as "no limit known" and the client simply does not draw the "1 of 2" hint — never as "unlimited", which would have the interface contradicting the node. --- ## 5. Nothing gets stuck The requirement the operator stated first, and the one worth over-engineering. Two directions, and both must be closed: - **a slot the node never gets back** — the node fills up and everyone queues forever; - **a transfer the client shows as waiting when the node has forgotten it** — one person's widget lies while the node is idle. ### 5.1 Every way a slot can be lost | How it ends | What reclaims the slot | When | |---|---|---| | Transfer completes | Client sends `transfer_close{done}` | Immediately | | …and the client forgets to | **Node auto-closes on serving the last chunk** — it knows `entry.size` and `CHUNK_SIZE`, so it knows the final index | Immediately | | User cancels | `transfer_close{cancelled}` | Immediately | | User pauses | `transfer_close{paused}` | Immediately | | Tab closed, browser quit, app killed | `shutdown_tasks()` / `_unregister_peer()` releases every lease of the session | On WebRTC `connectionstatechange` — the same hook that already ends streams | | Network drops, no clean close | Idle reclaim: no `file_req` under this `tr` for 120 s | ≤ 120 s | | Client granted a slot and never used it | Acceptance deadline | 30 s | | …and it *is* using it, but nothing said so | **`touch()` on every `file_req` carrying the lease.** Missing from the first build: the pool had the method, the tests covered it, and no caller existed — so `used` stayed False for every download and the acceptance deadline revoked each grant at 30 s while the file transferred at 20 MB/s | — | | A grant revoked, requeued, granted again, revoked again | **Bounded at three misses, then the lease is closed.** Also missing at first: the requeue was a permanent cycle and the node logged the same reclaim every 30 s until it restarted | — | | Client crashes mid-zip between two files | Idle reclaim (the gap between files is milliseconds) | ≤ 120 s | | Node restarts | Leases are in memory only | Immediately | | Group detached / member revoked mid-transfer | The existing revocation path drops the sessions, which releases their leases | Immediately | The first two rows are the answer to *"a user must not stay blocked once their downloads have been done"*: the client says so, **and** the node concludes it independently. Neither is trusted alone. ### 5.2 Every way a client can be left waiting | Failure | What fixes it | |---|---| | Node granted a slot, the push was lost | Client watchdog: no state change for 60 s while `queued` → re-send `transfer_open` (idempotent on `tr`) | | Reconnect: session gone, leases gone | `_onReconnected` (already in `transport.js`) re-opens the lease for every running/queued transfer, with `from_chunk` | | Node forgot the lease (idle reclaim during a stall) | Node pushes `transfer_state{state:"revoked"}`; the widget shows **interrupted — resume**, not a silent hang | | Node at cap and everyone is idle | Cannot happen if §5.1 holds; if it does, `GET /api/transfers` shows the leases and the operator can force-release one. Ship that endpoint | | Queue entry for a member who left the group | Pump skips entries whose session is closed and drops them | ### 5.3 The invariant, and how it is checked > **A new grant is never made past a cap, and every queue entry names a lease > that exists. A lease is created in exactly one function and destroyed in > exactly one function.** > > Not "granted leases never exceed the cap" — that was the first wording and the > property test rejected it within a second of being written. Lowering a cap > never interrupts a running transfer, so the count legitimately sits above the > new value until those finish. The invariant is about what may be *handed out*, > not about what is held. Two tests, not one: - a unit test per row of both tables above; - a **randomised property test** — a few thousand random sequences of open, close, drop, reconnect, pause, resume and cap changes, asserting after each step that the counter equals the number of live leases and that no queue entry references a dead session. The stuck-slot bug is a race by nature, and "it works now" is not evidence against a race (`window_leak.mjs` is the precedent in this repo). --- ## 6. Pause and resume Cancelling already works client-side (`transfers.cancel`) and needs only to send `transfer_close`. Pausing is new, and **whether it can be resumed depends entirely on where the file is being written**. The interface must offer only what the target can actually do — a pause button that quietly restarts the download from zero is worse than no pause button. ### 6.1 By download target | Target | Platform | Pause in-session | Resume after reload | Notes | |---|---|---|---|---| | File System Access (`showSaveFilePicker` / granted folder) | Chrome, Edge | Yes — keep the `writable` open, stop asking for chunks | **Yes** — reopen with `createWritable({keepExistingData:true})` and `write({type:"write", position})` | The full story. `downloads.js` already keeps the directory handle in IndexedDB | | Electron native sink | Desktop app | Yes | **Yes**, once `main.js` gains a `save:resume` opening the file with `flags:'a'` and returning its current size | ~30 lines in `main.js` + `preload.js` (`main.js:797`, `preload.js:150`) | | Service-worker stream | Firefox, Safari | **No** | No | Not on this target — see §6.5, which says what would be needed to change that. **Hide the pause button**, keep cancel | | Blob in memory | Fallback | Yes (it is just an array) | No | Bounded by `BLOB_LIMIT` anyway | | **Upload** (any platform) | All | **Yes** | **Yes**, via the sealed probe | A `File` is seekable and the *node* keeps the position, so there is no local target to consult. Missed entirely when 7a shipped — pause was built around the download path and uploads were refused it by the same guard that protects a transfer which cannot re-acquire its slot | **Corrected while implementing 7a.** `platform.capabilities` is the wrong home for this: resumability is a property of the *target*, not of the platform. The same Chrome yields a pausable target from a granted folder and an unpausable one from the service worker, on the same page, for two files in the same batch. So each target declares `pausable` itself, `prepare` carries it into the store, and the widget renders from that record. The precedent to follow — and the mistake not to repeat — is `_openDownloadTarget`'s silent fallback chain (`CLAUDE.md`: *"a fallback chain reaches its floor silently"*). Whatever the target ends up being, the transfer records which tier it got, and the widget renders from that record. ### 6.2 Pause releases the slot Stated again because it is a design decision, not an implementation detail: a paused transfer holds nothing. Resuming rejoins the queue at the tail, and the widget says so (*"resuming — 2 ahead"*). Anything else lets one member close the node by pausing. ### 6.3 Resuming the transfer itself Trivially available, because `pipelinedDownload` is already indexed by chunk: resume is `nextRecv = Math.floor(bytesWritten / CHUNK_SIZE)` with the window refired from there. Two constraints: - **`bytesWritten` must be a multiple of `CHUNK_SIZE`.** It is, because writes are sequential and whole chunks — but a resumed file whose size is not a chunk multiple (a partial write interrupted by a crash) must be **truncated down** to the last whole chunk before resuming, never appended to. A silently corrupted download is worse than a failed one. - **The file must not have changed.** The entry id is its blake3 (`GroupIndex`), so the resume record stores `{fileId, groupId, size, bytesWritten, targetRef}`. A resume whose `fileId` is no longer in the index fails with "this file has changed on the node", which is the truth. ### 6.4 Resuming an upload The node already has the state (`self._uploads`, `.part` on disk). Make it usable and stop it leaking: - **Corrected while implementing.** `transfer_open` travels in clear, and the node identifies an upload by (member, directory, filename) — so asking there would put the operator's filenames on an unsealed message, which is precisely what sealing this path bought in MNP 2.0. The question is asked inside the seal instead, as an ordinary `file_upload` with no bytes and `UPLOAD_PROBE_INDEX` (-1); the node answers `resume_from` in the sealed ack, writing nothing and reserving nothing. A node that predates it refuses the index, which reads as "start from the beginning". - Move `self._uploads` from the session to the **group context**, keyed by `(user_id, rel_dir, filename)`, so a reconnect finds it. It is already authorised per member; the session was never the right owner. - **Reap orphans.** A `.part` older than 24 h with no live lease is deleted, on a timer and at startup. This is a pre-existing leak (§1) and this work is where it gets fixed. - The no-overwrite rule, the filename allowlist, the size cap and the chunk ordering are **unchanged** — draft v6 §2.1 names those four as what makes an upload safe, and none of them is touched here. --- ### 6.5 Firefox and Safari: OPFS, and what a pause would cost > **Corrected three times on 2026-09-08, ending here.** This section first said pause > was impossible on Firefox; then that OPFS was the missing streaming target. > **Both were wrong, and measurement is what settled it.** Firefox 154's OPFS > quota is exactly 10% of the volume's size — ~12.8 GB on a 128 GB disk, hit > *mid-download* — so OPFS cannot carry a film. The unbounded path is the > **service worker**, which was already there and merely unreliable; it was > fixed on 2026-09-08 (see `memory-audit-large-files.md` §6.5 and > `test_streamed_download_reliability.py`) — and then fixed again the same > evening, because "reliable" turned out not to include *staying alive*: an idle > service worker is terminated after about thirty seconds and a streaming > response does not count as activity, so every download longer than that lost > its reader mid-file (§12.2). What survives below is only the pause question, > for which OPFS is still the candidate, under its quota. Deferred by decision, so this section records **why**, because "Firefox cannot pause" is not true as stated and would be a bad thing to leave in a plan. **Why the current target cannot pause.** On the service-worker path (`downloads.js:openStreamedDownload`) the browser is already writing an HTTP response to its own download folder. Three separate things break, and only the first is about our code: - pausing means not writing to the stream, which leaves the browser's download stalled in its own download manager — visible to the user, outside our control, and cancellable by them in a way we never hear about; - **a service worker is terminated when idle.** Firefox and Chromium both kill one after tens of seconds with nothing to do. A pause that outlives that — which any pause waiting on a queue does — takes the stream down with it; - resume after a reload is genuinely impossible on this target. The response is committed; no API reopens a browser-managed download to append to it. **The alternative that does exist: OPFS.** The Origin Private File System (`navigator.storage.getDirectory()`) is a real, seekable, persistent filesystem, private to the origin, and it **is** implemented in Firefox and Safari — which is exactly what those two lack when they lack File System Access. A download would be written there, resumable and reload-proof like any local file, and handed to the person at the end: `handle.getFile()` returns a `File` backed by disk, so `URL.createObjectURL` on it is a link the browser copies disk-to-disk without loading gigabytes into the tab. So it is not a browser limitation. It is a **second download target to write and maintain**, and it costs: - a fourth tier in `_openDownloadTarget`'s already four-deep chain — the function whose silent floor cost this project a gigabyte of RAM per film once already (`CLAUDE.md`); - **quota.** OPFS is subject to storage quota and to eviction unless `navigator.storage.persist()` is granted. A 4 GB film may simply not fit, and finding that out at 90% is worse than not offering the path; - **double disk usage and a real wait at the end.** The file exists twice while the browser copies it out, and the copy is not instant for a large file; - **cleanup.** Abandoned partial downloads sit in OPFS invisibly, consuming the origin's quota, until something reaps them. That is a new janitor with its own failure mode; - **an API matrix that must be measured, not read.** The OPFS write path is not uniform: `createSyncAccessHandle()` (workers only) and `createWritable()` have had different availability across Firefox and Safari versions. This repo's standing rule applies — *model the environment, never the code under test* — so the answer comes from a probe in real browsers, not from a compatibility table. **Recommendation.** Not in this work. Revisit it as its own change once transfer slots are in production, and if it is picked up, it is worth more than pause alone: it would give Firefox and Safari resumable downloads, reload-proof, and retire the service-worker path and its iframe entirely. Ship it as a target first and let pause fall out of it — the reverse order is how a fourth tier gets bolted onto the chain in a hurry. ## 7. Settings ### 7.1 Node-wide — draft v6 §2.11's pattern, exactly Two new `[node]` keys, `max_concurrent_downloads` and `max_concurrent_uploads`, default 8, positive int, `_positive()`-validated like `max_concurrent_streams` (`config.py:332`). Then, mechanically: - `roster.py`: `SETTING_MAX_DOWNLOADS` / `SETTING_MAX_UPLOADS`, added to `node_settings()` (`roster.py:943`); - `ops.py`: entries in `get_node_settings` / `set_node_settings` (`ops.py:1290`, `ops.py:1309`), written to `roster.db` **and** `node.toml`; - **hot-swap through one helper.** `ops.set_node_settings` currently pokes `webrtc._stream_sem`, which does not exist (§1). Replace all three with `webrtc.set_capacity(streams=…, downloads=…, uploads=…)` on `WebRTCTransport`, which resizes the live pools and pumps the queues. One function, on the object that owns the state, with a test that changes a cap on a running transport and watches a queued transfer start; - the Node page grows two number fields beside "max concurrent streams" (`node-page.js:1042`) and the CLI grows `meshbay-node transfers show|set` — plus `meshbay-node transfers list`, which prints live leases and the queue from `GET /api/transfers`. ### 7.2 Per group, per member — `apps_enabled`'s pattern, exactly `OP_TRANSFER_LIMITS` in `adminop.py`, subject `"d=2,u=2"` so what the operator signs names the outcome. `_do_transfer_limits` + `_admin_exec_transfer_limits` in `webrtc_server.py`, stored by `ops.set_transfer_limits` in `roster.db` group settings, broadcast as `transfer_limits_ack` to the group's peers, surfaced in the group Settings tab as a section beside the scan settings. **Absent means the default (2), not unlimited.** Deliberately unlike `member_upload`'s "absent means allowed": a group that predates the setting and came back unlimited would leave the node-wide cap as the only control, which is the situation this work exists to end. Enabled by default, breaking, no opt-out switch — the same call as MNP 2.0's chat encryption, for the same reason. No release note and no operator prose: the project is in development and every node is a test node. Bounds: 1–32. Zero is not "unlimited" and is not accepted; a member who may not transfer at all is a member the operator revokes. --- ## 8. The client ### 8.1 `TransferStore` — three new states, one changed contract ``` ┌──────────┐ slot granted ┌──────────┐ start() ─────▶│ queued │────────────────▶│ running │ └──────────┘ └──────────┘ ▲ │ │ │ │ resume() │ │ cancel() pause() │ │ │ done │ ▼ ▼ │ ▼ ┌──────────┐ ┌────────┐ ┌──────┐ │ paused │◀──────────────│ paused │ │ done │ └──────────┘ └────────┘ └──────┘ │ │ cancel│ failed / cancelled ▼ ┌───────────┐ │ cancelled │ └───────────┘ ``` Statuses become `queued | running | paused | done | cancelled | failed`, and `run({signal, onProgress})` becomes `run({signal, onProgress, lease, resumeFrom})`: - `lease` — an object the store hands the runner: `await lease.acquire()` blocks until granted, `lease.release(reason)` in a `finally`. The runner never speaks to the transport about slots directly. One place opens a lease, one place closes it — the same rule as the node's. - **The slot is asked for after there is somewhere to write, and that ordering is load-bearing.** Asking first reads better — the widget could draw a row while the target is being chosen — and is wrong: opening a target takes tens of seconds of streamed-download timeouts, or as long as somebody leaves a Save As dialog open, and a grant not taken up in time is revoked. Tried the other way during the build: three downloads started, one arrived. A test pins the order now. - The store gains `pause(id)`, `resume(id)`, `pauseAll()`, and a persisted resume record for the tiers that support it. - `_maybeRelease(transport)` must treat **paused and queued as busy** (`transfers.js`'s `_busy`). A paused transfer whose transport was closed because the group page went away can never resume — that is a one-line regression waiting to happen, and it has a test. ### 8.2 The widget `TransferWidget` (`app.js:141`) is a flat list of rows with a progress bar. It becomes a small panel that can answer, at a glance, *what is happening and what is my fault*: ``` ┌─ Transfers ───────────────────── 2 running · 3 waiting ─┐ │ ⏸ Pause all Clear finished │ ├──────────────────────────────────────────────────────────┤ │ ⬇ Some Saga (2019).mkv ⏸ ✕ │ │ ███████████████████░░░░░░░░░░ 62% │ │ 2.4 GB / 3.9 GB · 11.2 MB/s · 2 min left │ ├──────────────────────────────────────────────────────────┤ │ ⬆ holiday-photos.zip ⏸ ✕ │ │ ████████░░░░░░░░░░░░░░░░░░░░ 27% │ │ 118 MB / 430 MB · 3.1 MB/s · 1 min left │ ├──────────────────────────────────────────────────────────┤ │ ⬇ A Different Show S02E04.mkv ✕ │ │ ┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈┈ waiting · 1 ahead │ │ Your 2 slots are busy │ ├──────────────────────────────────────────────────────────┤ │ ⬇ render.tar ▶ ✕ │ │ ██████░░░░░░░░░░░░░░░░░░░░░░ paused at 21% │ ├──────────────────────────────────────────────────────────┤ │ ⬇ notes.pdf Finished ↗ │ └──────────────────────────────────────────────────────────┘ ``` Concretely: - **A summary in the header** — `2 running · 3 waiting` — and the nav badge counts running *and* waiting, so a queued transfer is never invisible. - **Waiting rows have a striped, indeterminate bar** (a CSS gradient animation, no JS), visibly different from a stalled progress bar. `prefers-reduced-motion` turns the animation off, not the state. - **The reason for waiting is written out**, from the node's own counters: *"Your 2 slots are busy"* vs *"The node is at capacity — 1 ahead"*. Someone should be able to tell their own limit from the operator's without asking. - **Pause / resume / cancel as icon buttons**, with real `aria-label`s and `title`s; pause is **absent, not disabled**, where the target cannot resume (§6.1) — and the row carries a quiet hint saying why, once. - **ETA** from the existing windowed speed, hidden below 5 s of samples so it does not flicker nonsense at the start. - **Ordering**: running, then waiting (queue order), then paused, then finished. A row does not jump when its neighbour finishes — group headers rather than a re-sort on every emit. - **Live region**: state transitions announce once (`aria-live="polite"`), progress does not. - Colours from the existing CSS variables, so light and dark both work without a second palette. New rules go in `style.css` beside `.transfer-*`, measured with `tests/harness/layout_probe.py` at 320 px — the transfers panel is *precisely* what that harness was written for, after it hung 138 px off the left of a phone. ### 8.3 i18n New keys: `transfers.waiting`, `transfers.waiting_ahead`, `transfers.paused`, `transfers.pause`, `transfers.resume`, `transfers.pause_all`, `transfers.your_slots`, `transfers.node_busy`, `transfers.eta`, `transfers.interrupted`, `transfers.no_pause_here`, `transfers.running_count`, plus the settings labels. `en.js` is the source; `test_locales.py` holds the other nine to its key set and will fail until all ten are translated. --- ## 9. Regressions to watch Ranked by how quietly they would ship. 1. **Browsing degraded by a busy node — the operator's stated requirement (§3.4).** `_try_serve_thumbnail` shares `file_req` with real files, and the Photos viewer and the Files preview modal fetch whole index entries through `pipelinedDownload` just as a download does. Four tests, not one: with both pools full and a queue behind them, (a) a Videos poster grid still fills, (b) a Music album grid still fills, (c) a photo opens full size, (d) a PDF preview opens. Each must also fail with the exemption removed. 2. **Video streaming charged twice.** `stream_req` must not touch the download pool. Test: fill the download pool, start a film, assert it plays. 3. **Chat attachments — decided: they take an upload slot** like any other upload, through the same `_do_file_upload` path. The regression is therefore in the *interface*, not the rule: a paperclip whose file silently sits in a queue is a chat that looks broken. `chat-app.js`'s composer must render the same waiting state as the transfers panel — the attachment row says "waiting, 1 ahead", not nothing — and a queued attachment must not block the composer, which is the freeze already recorded in `CLAUDE.md` for chat sends. Test it in `chat_send_probe.py`, which drives composer and transport together for exactly this class of seam. 4. **`_maybeRelease` closing a transport with paused transfers on it** (§8.1). 5. **Zip downloads.** One lease for the whole job. Getting this wrong means a 40-file zip takes 40 slots and deadlocks against its own cap — a deadlock, not a slowdown, because the job cannot finish until it holds them all. 6. **Search across groups** (`search-page.js`, `getTransport`) opens transports to several nodes; each node caps independently. Correct, but the widget must not present another node's queue as this one's. 7. **`_free_name` and resumed uploads.** A resumed upload must reuse `state["stored_name"]`, not draw a new one — otherwise a reconnect produces `file (2).mkv` next to a half-written `file.mkv.part`. 8. **The 60 s cap in `_do_file_request`'s backpressure loop** interacts with a paused reader: a paused transfer stops requesting, so nothing is waiting — but check that a *slow* reader is not now reclaimed as idle. The idle timer is reset on each `file_req` under the lease, not on each chunk *sent*. 9. **`test_security_regressions.py:786`** asserts `_transcode_sem` is in the source. Renaming it during the `set_capacity` refactor will fail that test — which is the test doing its job. Update it deliberately. 10. **The one this list did not anticipate, and the one that happened:** none of the eight defects the live pass found are in this list, because every entry here is about *slots* and seven of the eight were about the download path underneath them (§12.2). A regression list written from the change being made will not cover the ground the change stands on. Before the flag day, walk the download path itself — every write target, on every browser — rather than the diff. --- ## 10. Tests Following the repo's rule: measure the environment, never model the code under test. **Node (`packages/meshbay-node/tests/`)** - `test_transfer_slots.py` — real `WebRTCPeerSession` objects against a fake DataChannel (the shape `test_webrtc_transport.py` already uses): node cap, per-member cap across two sessions of one account, queue order, head-of-line skip, grant deadline, idle reclaim, auto-close on the last chunk, release on `shutdown_tasks`, no double-charge on a re-`transfer_open` with the same `tr`. - `test_transfer_invariant.py` — the randomised property test of §5.3. - `test_transfer_exemptions.py` — the §3.4 requirement, as behaviour: posters, covers, thumbnails, a full-size photo and a document preview all served with both pools full and a queue waiting; the 2-in-flight bound refusing a third concurrent leaseless entry; and a source-reading check that the only callers opening a lease are the three that call `transfers.start()`. - `test_navigation_not_leased.py` — `index_sync`, `index_delta`, `media_meta_req`, `music_meta_req` and directory listing answered normally with both pools full. Cheap, and it fails the day someone routes a listing through `file_req`. - `test_upload_resume.py` — reconnect mid-upload resumes at `next_index`; orphan `.part` reaped; `stored_name` preserved. - `test_node_settings.py` extension — `set_capacity` actually resizes a live pool (the test `ops.py:1362` never had). **Hub / SPA (`packages/meshbay-hub/tests/`)** - `test_transfers.py` extension, under Node as today: queued → running, pause releases, resume re-queues, cancel while queued, `_busy` counts paused. - `tests/harness/transfers_probe.py` — mounts the real `TransferWidget` in Chrome against a stubbed store and reads the rendered states back. The precedent is `chat_send_probe.py`: the seam between store and widget is where this will break, and neither source shows it. - `test_layout_measured.py` extension — the panel at 320 px with a waiting row. - `test_locales.py` — passes only when all ten catalogues have the new keys. **Live — `packages/meshbay-node/tests/transfer_probe.py`, in the repo** Not in `QE/` as this section first said: it found four defects nothing else could, and `QE/` is not versioned, so it lived on one machine. It is not collected by pytest (the filename does not match `test_*.py`) and still needs `QE/deploy/e2e.py` and `demo.env` at run time, which it locates and explains rather than importing blindly. `--want N` measures the cap, `--pull N [--parallel]` downloads real files to completion on one connection, and `--operator` covers the two things only the operator's CLI can answer: that a cap raised live starts what was waiting, and that a vanished peer's slots are back before anyone asks. - ~~`QE/deploy/transfer_probe.py`~~ — opens N real MNP sessions as two accounts, starts more transfers than the caps allow, and asserts the observed concurrency and the order they complete in. The equivalent of `stream_probe.py`, and the only thing that answers "is it the node or the browser" in one run. - A two-browser manual pass: cap 2, start 5, watch the queue drain; pause one, reload the page, resume it; kill a tab mid-transfer and watch the slot come back within 120 s. - `QE/migration/reap_orphan_parts.py` — one-shot, node stopped, for the `.part` files already on deployed nodes (§6.4). Per the QE rule in `CLAUDE.md`: any test against meshbay.org opens the UFW port, tests, and closes the port and kills the processes **in the same block**. --- ## 11. Order of work Each stage is shippable and testable on its own; nothing before stage 5 changes what a member sees. | # | Stage | Contents | |---|---|---| | 1 | ✅ **Capacity, fixed** | `set_capacity()` on `WebRTCTransport`, replacing the dead `_stream_sem` poke. No new features. Ships the §1 bug fix alone, where it can be verified alone | | 2 | ✅ **Leases, node-side** | Lease table, pools, queue, pump, every reclaim path in §5.1, `GET /api/transfers`, the DEBUG line. No client uses it yet; the node grants everything because no client asks | | 3 | ✅ **Settings** | `[node]` keys, roster overrides, `OP_TRANSFER_LIMITS`, Node page fields, CLI verbs, ack field | | 4 | ✅ **MNP 3.0** | Version bump on both sides, `MNP_MIN_SUPPORTED` at 3.0, **§3.4.1's bound on leaseless reads** (which did not exist), the desktop client's version gate, every package aligned on 0.13.0 | | 5 | ✅ **Client leases** | `TransferStore` states and the lease contract, `file_req`/`file_upload` carrying `tr`, reconnect re-open, watchdog | | 6 | ✅ **Widget** | The panel of §8.2, i18n, layout measured | | 7a | ✅ **Pause / resume, in session** | Per-target `pausable`, the slot released on pause and re-asked on resume, `fromChunk` in the pipeline, the widget's pause button | | 7b | ⬜ **Resume across a reload** | The persisted resume record, `save:resume` in Electron, `createWritable({keepExistingData:true})` with a position, truncate-to-chunk | | 8a | ✅ **Upload state + reaping** | `uploads.py`, state in the group context keyed by member, the `.part` janitor | | 8b | ✅ **Upload resume + pause** | The sealed probe chunk, `resume_from`, uploads asking for their own slot, `touch()` on an upload chunk | | 9 | ✅ **Live pass** — done out of order, and §12 is its report. | `transfer_probe.py`, two browsers, the manual list | | — | *Deferred, separate change* | **OPFS as a download target** for Firefox and Safari (§6.5) — resumable downloads there, and the retirement of the service-worker path. Not part of this work | Stages 1–3, 5 and 6 have landed. Stage 9 was run **before** stage 4 rather than after, deliberately: while nothing is enforced, a defect in the machinery costs nothing, and §12 is what that bought. Stage 4 is a flag day and should land only on the evidence §12.3 describes. --- ## 12. What the live pass found, and what it changes Step 9 was meant to confirm the machinery. It found **eight defects**, seven of which no test in the repo could reach, and three of which are in code the sections above describe as settled. Four more (§12.4 to §12.7) came out of fixing those eight, and the last of them — a hard reload silently disabling the only unbounded write path Firefox has — was the one actually being reported all along. They are listed here because their pattern matters more than any one of them: every one needed a real browser, a real node and a real hub to show itself, and the last two were only reachable on Chrome. ### 12.1 The three in the lease machinery | | Found by | |---|---| | **`touch()` was never called.** The node ignored `tr` on `file_req`, so `used` stayed False for every download and each grant was revoked at 30 s mid-transfer | the node's own log | | **The requeue was a permanent cycle.** Revoked → queued → granted again → revoked, every 30 s, for as long as the daemon ran | the node's own log | | **`transfers show` reported the module defaults**, so `transfers set 2 2` answered "applied now" and the next line said 0/8 | typing the command | The first is the one to learn from. `TransferSlots.touch()` existed, was covered by its own test, and **had no caller**. The pool was right, the handlers were right, and the call between them was missing — so neither side's tests could see it, and the property test could not either: nothing drifted. **A seam is not tested by testing both sides of it.** ### 12.2 The five in the download path itself None of these are about slots, and all of them would have become "this download is impossible" the moment leases were compulsory: - **three headers** decide whether the page may frame its own `/_mbdl/` URL — `frame-src`, `frame-ancestors`, `X-Frame-Options` — and all three were wrong. Fixed one at a time over an afternoon; all three were visible in a single `curl -I` against the deployed hub; - **a service worker with no event for ~30 s is terminated**, and a streaming response does not count as activity. The reader vanished mid-file and `writable.write()` never resolved *and never rejected*: no error, no log, a progress bar that stopped, and a node that stayed healthy throughout; - **`encodeURIComponent` leaves `'` alone** and `'` is RFC 5987's delimiter, so a 449 MB film arrived complete and correct under the name `mtsshk9w-ohqty535`; - **a browser grants one file picker per user gesture**, so downloading three files at once failed two of them with a message about gestures; - **two silent returns** meant a click with no connection produced nothing at all — no transfer, no icon, no message. ### 12.3 What this changes about step 4 Step 4 makes leases compulsory and refuses 2.x peers at the handshake. It is the only irreversible step in this plan. Every defect in §12.1 and §12.2 was invisible to 1169 node and 814 hub tests, and was found by a person clicking Download and pasting a console. Three of them were introduced *by this work* on the day it was written. Making a mechanism compulsory is a bet that it is well understood, and the evidence of one afternoon is that it was not. **The recommendation was to wait for evidence rather than for time**: a week of ordinary use with no freeze and no lost transfer, or the equivalent in deliberate runs of `transfer_probe.py --pull 3 --parallel` and `--operator`. **Overruled by the operator on 2026-09-09, and the reasoning is better than the recommendation it replaced.** The concern was that a flag day removes the fallback exactly when defects are still being found; the answer is that the fallback only has value while more work is coming down the same path, and none is — the next work is a music application with playlists, which touches none of this. The prerequisite this section named was satisfied first: `GET /v1/hub/version` already carried `client.minimum`, and the desktop client now reads it before connecting instead of meeting a handshake refusal it has no vocabulary for. Nothing about §4's reasoning has changed. What changed is the confidence that the thing being made compulsory works. ### 12.4 The silent row — fixed **The transfers panel showed nothing while the target was being opened.** The row was created after `_openDownloadTarget` returned, which can take tens of seconds, so three clicks produced no icon at all and then several rows at once. It was "fixed" during the build by taking the lease first, which caused §12.1's `not_taken_up` symptom and was reverted — on a wrong diagnosis, as it turns out: the revocations were the missing `touch()`. The revert was right anyway (§8.1). The shape that worked is a `prepare` step in the store, distinct from `run`: the row appears at the click, the target is opened, and the lease is asked for last. The reservation about `showSaveFilePicker` needing a user gesture turned out to be real, but not in the way expected — see §12.5. ### 12.5 One dialog per file, and three downloads frozen behind it Reported from Chrome, after §12.4 shipped: selecting four files produced a Save As dialog for the first, then — once that file had finished — a dialog for the second, while the last two timed out. On a later attempt the three remaining transfers simply froze. Three facts explain it, and only the third was a surprise: 1. `for (const entry of selected) await downloadFile(entry)` used to serialise the target openings **by accident**. Moving the opening into `prepare` removed the accident and four pickers raced. Fixed by `_openTargetInTurn`, an explicit queue on the targets — never on the rows, so every download still appears at the click. 2. A browser grants one file picker per user gesture, and selecting four files is one gesture. The code already recovered from the `SecurityError` Chrome throws for a picker with no gesture behind it, by streaming instead. 3. **Chrome does not throw.** It shows the dialog anyway and waits for a human. So the recovery in (2) was never reached, and the queue from (1) turned the unanswered dialog into a head-of-line block: the third and fourth downloads were not frozen, they were waiting correctly for a dialog nobody had been shown yet. From the panel that is indistinguishable from a freeze. The fix is to stop asking. `_openTargetInTurn` marks everything that has to wait its turn as `batched`, and a batched opening prefers the streamed path whatever the download mode says. The first file of a batch — the one that actually holds the gesture — still gets its dialog, so the preference is honoured where it can be. For the rest there is no gesture left to spend, so nothing is lost by streaming: the file still lands on disk, in the browser's own download folder, written as it arrives. Only the choice of folder goes, and it was not on offer. If the worker does not answer, a batched download falls back to the dialog rather than failing — asking is better than losing the file (§ "a preference must not cost a capability", which applies to the fix as much as to the bug). Pinned by `test_targets_are_opened_one_at_a_time` (the queue, and that the first opening is the only unbatched one) and by four cases in `test_memory_ceiling.py` (the branch itself, both fallbacks, and that batching never pushes a large file into memory). **What that left, and what was measured.** With the queue unblocked all four files downloaded, but a dialog still appeared for each one — with MeshBay's own setting on "save automatically", which should never prompt. That points at the streamed path failing, so it was measured rather than reasoned about: a real Chrome 152 driven over CDP against the deployed hub, running the actual flow (a `TransformStream` posted to the worker, a hidden iframe on `/_mbdl/`, a wait for `mbdl-serving`). It is served in 2–3 ms on a normal load, after a hard reload, and twice in the same document. Ctrl+F5 does leave the page uncontrolled — `navigator.serviceWorker.controller` is null and no `controllerchange` arrives — but the `mbdl-claim` recovery already in `_claimController` gets control back inside 3 s. Both probes are in `QE/`-style scratch scripts, not in the suite: they need the deployed hub. **And the queue itself became the next defect.** Serialising the openings was new in the same commit, and on Firefox it regressed what had always worked: four downloads that opened their targets at the same time began waiting on the slowest, and all four sat at "preparing" — the node journal showing `d=0/8(q0) u=0/8(q0)`, not one transfer opened, so the block was entirely in the client before any slot was asked for. Measured on Firefox 154 against the deployed hub: `register` and `ready` return instantly, the page is controlled, and four serialised openings are served in 5–18 ms — so the streamed path was not the delay either. The queue was. Two bounds fix it, and both are narrowings of the queue rather than of any capability. Only an opening that could actually show a dialog joins the queue, which on Firefox and Safari — no `showSaveFilePicker` at all — means none of them do, restoring exactly the previous behaviour. And no opening waits behind another for longer than `TARGET_QUEUE_BUDGET_MS` (90 s), because `_targetQueue` is never reset and an opening that never settles would otherwise leave the page unable to start a download again until it is reloaded. Releasing early is safe: whatever is ahead is still the only unbatched opening, so the released one takes the streamed path and opens no second dialog. The general shape, third time in this section: **a queue is a way to convert one slow participant into several stuck ones.** Every queue added here needs to say what it protects and refuse everything it does not protect, and needs a bound. ### 12.6 The wait with no deadline Bounding the queue was not enough: Firefox still showed four rows at "preparing", with the queue bypassed, so each opening was hanging on its own. `_claimController` had two waits with no deadline of any kind — `navigator.serviceWorker.register()` and `navigator.serviceWorker.ready` — while `SW_CONTROL_BUDGET_MS` bounded only the wait that comes *after* them. And `_swPromise` is shared, so a single unsettled one of those left every download on the page waiting on the same promise, for the life of the page. Measured on Firefox 154, on a local `127.0.0.1` site so no hub was involved: | worker | `register()` | `ready` | | --- | --- | --- | | that installs | 8 ms | 0 ms | | **whose install handler rejects** | **7 ms** | **never settles** (still pending at 10 s) | That is the whole mechanism. `register()` resolves as soon as the registration object exists — carrying nothing but an *installing* worker — and `ready` is what waits for an active one. A worker that cannot install therefore produces a registration that looks fine and a `ready` that never comes. Every wait in `_claimController` is now inside one budget, with two carve-outs that exist so a deadline never costs a capability. A `ready` that times out while `registration.active` is set is not fatal — `ready` may be waiting on a newer worker that cannot install while an older one serves perfectly well. And the `mbdl-claim` recovery keeps its own budget outside the deadline, because giving up there would cost Firefox the only unbounded way it has to write a download to disk. **A deadline alone would still not have been a fix**, only a better-explained failure: a registration stuck with nothing but an installing worker does not heal, and every later visit finds the same one. So when `ready` times out with no active worker, the registration is discarded (`unregister()`) and asked for once more with a fresh budget. The page repairs itself instead of needing developer tools. Pinned by four cases in `test_streamed_download_reliability.py`, each checked against the unfixed source: a worker that never installs, a registration that never answers, a stuck `ready` that must not throw away a working worker, and the discard-and-retry. So the streamed path is not what is prompting. Exactly three things can open a dialog per file on Chrome in automatic mode: our `showSaveFilePicker`, the folder permission bubble from `ensurePermission` (only when a folder was granted and the permission lapsed, and it is not a Save As), and **Chrome's own "always ask where to save each file"** in `chrome://settings/downloads` — which applies to the worker's response because it carries `Content-Disposition: attachment`, and which no download mode of ours can override. The `console.info` added with this fix distinguishes them in one line: it is written before every dialog we open, so a dialog with nothing in the console is the browser's. ### 12.7 The one none of the above explained: a hard reload Every fix in §12.5 and §12.6 was real, and none of them was the defect being reported. Downloads on Firefox failed with "the worker did not answer the download within 15s", every time, for the operator, while the *same profile* driven from this machine — headless and headed, on their display, with their files, through the real interface — succeeded every time. The operator's own sequence found it, and it is worth writing down verbatim because no automated run could have produced it: a freshly started browser downloaded four files out of four, twice; one **Ctrl+F5** and every attempt afterwards failed; restart, fine again; Ctrl+F5 before any attempt and the very first one failed. **A document fetched by a hard reload is loaded with the service worker bypassed.** It can still be claimed afterwards — so `navigator.serviceWorker.controller` comes back and every check in `_claimController` passes — but the navigations that document starts keep missing the worker, and the hidden iframe a streamed download needs *is* a navigation. On Firefox and Safari that is the only way to write a file too large to hold in memory, so the download cannot happen at all, for the life of that page. The same behaviour had already been measured on Chrome in §12.6's work and its significance was missed. The hard reloads were on this author's instruction, after each deployment. The SPA's HTML is served `no-store`, so an ordinary reload has always picked up a new build and Ctrl+F5 was never needed for anything. **Why the measurements could not find it.** WebDriver cannot perform a hard reload — the key event goes to the content, not the browser chrome — so every automated run tested, over and over, the one case that works. A fix is validated by tests; a diagnosis is not always validated by automation, and the operator's manual sequence should have been asked for hours earlier. The remedy is to stop inferring servability from control and to ask instead. At boot the client opens a four-byte stream and a hidden iframe, exactly as a real download would, and tears both down. If the worker does not answer, the page reloads **once**, ordinarily, which puts it back under the worker; the flag is in `sessionStorage` because it has to survive the reload it triggers and must stop rather than loop. Two delays then had to go, because the repair first landed about thirty seconds late — long enough to click and watch four rows hang: the claim is asked for *before* the control budget rather than after it (a page uncontrolled beside an active worker is never claimed on its own), and a download that starts while the self-test is running waits for it instead of racing it. And the refusal message no longer tells a Firefox user to switch to Chrome for a state an ordinary reload undoes. All ten catalogues say to reload first. --- ## 13. Decisions — all settled (operator, 2026-09-08) Nothing in this plan is waiting on an answer. | # | Decision | Consequence | |---|---|---| | 0 | **Browsing is never subject to a slot** — posters, thumbnails, covers, previews, listings | §3.4, structural: a transfer is what the widget shows | | 1 | **MNP 3.0, hard refusal of 2.x at the handshake** | No leaseless branch survives anywhere. Flag day, stage 4 | | 2 | **Per-member default of 2 applies to existing groups** | No release note, no migration prose — the project is in development and every node is a test node | | 3 | **Chat attachments take an upload slot like any other upload** | §9.3: the composer must show the waiting state, or the paperclip looks broken | | 4 | **Pause on the service-worker path (Firefox, Safari) is deferred, not abandoned** | §6.5 says what it would actually take. It is cost, not impossibility | | 5 | **The two big caps stay node-wide** | The resource is the machine's. A per-group variant can be added later with no protocol change |