diff options
| author | Christophe Besson <cbesson@gmail.com> | 2026-08-26 15:09:20 +0200 |
|---|---|---|
| committer | Christophe Besson <cbesson@gmail.com> | 2026-08-26 15:09:20 +0200 |
| commit | db7f81fd08742847f3ebea061c75530b7b31b934 (patch) | |
| tree | 14b308686860eea0a9c12c47c39f8a68a1bba376 /packages/meshbay-hub/src/meshbay_hub/static/file-utils.js | |
| parent | 7126fd3c265ba75d77b449bfd0f83f5f3e584b74 (diff) | |
| parent | 59d9f50bf41b2b38b7c95f8da52b698dff923c2d (diff) | |
| download | meshbay-db7f81fd08742847f3ebea061c75530b7b31b934.tar.gz | |
Merge branch 'debug/webrtc-lock-resume'
WebRTC transport dies silently after an extended mobile screen lock
(confirmed live via client-side trace + node logs): ICE goes
disconnected -> failed within ~10s of each other on both ends, but the
DataChannel's readyState stays "open" throughout, so nothing failed fast —
every request just sat out its own timeout, matching the reported symptom
(poster spinners, blocked chat, dead new streams, stuck music).
- Automatic reconnect on WebRTC "failed": capped exponential backoff,
redoes the full signaling handshake, wakes immediately on
visibilitychange instead of waiting out a throttled backoff timer.
- Fixed two real bugs the reconnect work exposed: the signaling POST to
the hub kept using the token captured at construction, never the fresh
one fetched per reconnect attempt (401 loop, no possible recovery); and
connect() re-armed a diagnostic listener/interval on every attempt
without disposing of the previous one.
- pipelinedDownload retries a lost chunk instead of aborting the whole
transfer — covers Files downloads, video poster/thumbnail fetches, and
music-player.js's blob-based track download.
- music-player.js: don't throw "Transport not connected" while a
reconnect is already landing (waitForReconnect); prefetch depth now
adapts to network type (5 tracks ahead on Wi-Fi, 3 on cellular or
unrecognized — Firefox/Safari included, where the detection API is
simply absent).
- video-player.js: onReconnected reissues the existing seek-to-current-time
path, so a mid-stream reconnect looks like an ordinary seek rather than
a dead player; holds a Screen Wake Lock unconditionally while open.
- New opt-in (off by default) user preference: keep the screen on during
audio playback, for whoever wants to trade battery for sidestepping the
screen-lock gap entirely — off by default because the ordinary
expectation (matching Spotify/Deezer) is that the phone locks on its
own while listening.
- hub: /app and / now serve Cache-Control: no-store — the SPA shell had no
cache header at all, so a browser that cached it heuristically could
keep re-serving an old build (old ASSET_V, old JS) through any number of
reloads or pull-to-refreshes.
Verified against real production use across many rounds (demo groups,
actual mobile screen-lock testing) rather than synthetic reproduction
alone. 430 hub tests + 650 node tests passing throughout.
Diffstat (limited to 'packages/meshbay-hub/src/meshbay_hub/static/file-utils.js')
| -rw-r--r-- | packages/meshbay-hub/src/meshbay_hub/static/file-utils.js | 37 |
1 files changed, 36 insertions, 1 deletions
diff --git a/packages/meshbay-hub/src/meshbay_hub/static/file-utils.js b/packages/meshbay-hub/src/meshbay_hub/static/file-utils.js index b44d105..ba76ac9 100644 --- a/packages/meshbay-hub/src/meshbay_hub/static/file-utils.js +++ b/packages/meshbay-hub/src/meshbay_hub/static/file-utils.js @@ -117,6 +117,41 @@ function _b64ToU8(b64) { return arr; } +// A dead transport (screen-lock WebRTC failure, see transport.js's +// _reconnectLoop) surfaces here as a rejected fetchChunk — TransportLostError +// when the pending request was killed outright, a plain timeout if it was +// still waiting when this ran. Either way the chunk itself was never the +// problem, and the file already on disk (writable has real bytes in it by +// now) is worth more than an all-or-nothing download: retry the same chunk +// instead of letting one bad moment abort the whole transfer. Each retry +// re-enters transport.fetchChunk, whose own _sendAndWait waits out an +// in-flight reconnect before trying again, so this loop is mostly just +// giving that reconnect the time and the attempts to land. +const CHUNK_RETRY_ATTEMPTS = 6; +const CHUNK_RETRY_DELAY_MS = 1500; + +function _isRetryableTransportError(err) { + return err.name === 'TransportLostError' + || err.message === 'Response timeout' + || (err.message || '').startsWith('DataChannel not open'); +} + +async function _fetchChunkResilient(transport, fileId, index) { + let lastErr; + for (let attempt = 0; attempt < CHUNK_RETRY_ATTEMPTS; attempt++) { + try { + return await transport.fetchChunk(fileId, index); + } catch (err) { + if (!_isRetryableTransportError(err)) throw err; + lastErr = err; + if (attempt < CHUNK_RETRY_ATTEMPTS - 1) { + await new Promise((r) => setTimeout(r, CHUNK_RETRY_DELAY_MS)); + } + } + } + throw lastErr; +} + async function pipelinedDownload(transport, gekKey, fileId, totalChunks, onChunk, writable, signal) { const results = writable ? null : new Array(totalChunks); @@ -125,7 +160,7 @@ async function pipelinedDownload(transport, gekKey, fileId, totalChunks, onChunk const fire = () => { while (nextSend < totalChunks && nextSend - nextRecv < PIPELINE_WINDOW) { - inflight[nextSend] = transport.fetchChunk(fileId, nextSend); + inflight[nextSend] = _fetchChunkResilient(transport, fileId, nextSend); nextSend++; } }; |