summaryrefslogtreecommitdiffstats
path: root/docs
diff options
context:
space:
mode:
Diffstat (limited to 'docs')
-rw-r--r--docs/MESHBAY_DESIGN.md21
-rw-r--r--docs/MESHBAY_NODE_PROTOCOL.md2
2 files changed, 14 insertions, 9 deletions
diff --git a/docs/MESHBAY_DESIGN.md b/docs/MESHBAY_DESIGN.md
index 7ce1168..8f91a87 100644
--- a/docs/MESHBAY_DESIGN.md
+++ b/docs/MESHBAY_DESIGN.md
@@ -2052,14 +2052,19 @@ MSE string) fed to a source buffer, with the node holding one slot per viewer.
from the end and echoed back; the client supplies the timestamp offset, because
copying timestamps does not preserve position.
- **`stream_init.start` names where the picture begins, never where the viewer
- dragged.** Copied video can only begin on a keyframe, so on that path the node
- resolves the request to the keyframe at or before it, seeks to that, and reports
- that. The client builds `SourceBuffer.timestampOffset` out of this number, and a
- subtitle cue carries the source's own absolute time: one GOP of disagreement
- between the two puts every line on screen before it is spoken. The look-up is a
- bounded read of the thirty seconds before the request, measured at 0.12–0.51 s
- on a real title, and only the copy path needs it — re-encoded video begins
- exactly where it is asked to.
+ dragged**, and on the copy path that position is **measured, never predicted**.
+ The client builds `SourceBuffer.timestampOffset` out of this number and a
+ subtitle cue carries the source's own absolute time, so every second of
+ disagreement puts a line on screen a second away from the voice saying it.
+ Reading the key frames and taking the last one at or before the request answers
+ a different question twice over: Matroska's Cues index only some keyframes, so
+ an index seek backs off to an indexed one that can be much earlier, and the
+ landing point moves with **which streams are mapped**, because the container is
+ positioned where every mapped stream has data — one real seek answered 4909.863 s
+ from the frame list and delivered 4907.236 s. So the node runs the same seek with
+ the same mapping, copies one frame under `-copyts`, and reads the answer back:
+ 0.06–0.07 s, cheaper than the scan it replaced. Only the copy path needs it —
+ re-encoded video begins exactly where it is asked to.
- **A seek trims what it can, and what it can differs per stream — which is a
desync, not an inconvenience.** Accurate seeking cannot trim copied video, which
must begin on a keyframe, but it does trim re-encoded audio to the exact request.
diff --git a/docs/MESHBAY_NODE_PROTOCOL.md b/docs/MESHBAY_NODE_PROTOCOL.md
index 2aede26..a53d494 100644
--- a/docs/MESHBAY_NODE_PROTOCOL.md
+++ b/docs/MESHBAY_NODE_PROTOCOL.md
@@ -1671,7 +1671,7 @@ array while MediaSource consumes it a segment at a time.
| Concurrent transcodes | 8 node-wide, semaphore on the transport context |
| Seeking | a new `stream_req` with `start`; the previous stream is retired first, ffmpeg respawned with `-ss` |
| Accurate seek | **off when the video is copied, on when it is re-encoded.** Copied video has to begin on a keyframe and cannot be trimmed to the request; re-encoded audio can, and is. Leaving both at the default put a whole GOP of silence at the head of every seek and left sound and picture a GOP apart — with correct timestamps throughout, so nothing downstream could detect it |
-| `start` in `stream_init` | the value actually used, and on the copy path that is the **keyframe at or before the request**, resolved by a bounded look-up before ffmpeg is spawned (0.12–0.51 s, measured). The client adds it back as `SourceBuffer.timestampOffset`; a subtitle cue carries the source's absolute time, so reporting the request instead would put every line on screen one GOP before it is spoken |
+| `start` in `stream_init` | the value actually used. On the copy path it is **measured** before the stream is served — the same seek with the same stream mapping, one frame copied under `-copyts`, 0.06–0.07 s — because an index seek lands on an indexed keyframe that the frame list cannot predict and that moves with which audio track is mapped. The client adds it back as `SourceBuffer.timestampOffset`; a subtitle cue carries the source's absolute time, so reporting anything else puts every line on screen away from the voice |
| `audio_tracks` in `stream_init` | every audio track: `i` (the **audio ordinal**, what `-map 0:a:<n>` takes, never the container stream index), `lang`, `title`, `codec`, `ch`. Empty for a file with no audio |
| `audio_track` in `stream_req` | which ordinal to map. Absent, out of range or malformed is the first track |
| `audio_track` in `stream_init` | the ordinal actually used, for the same reason `start` is reported: a list drawn before the file was replaced on disk can name a track that is no longer there, and the client must show what is playing rather than what it asked for. `null` when the file has no audio |