diff options
| author | Christophe Besson <cbesson@gmail.com> | 2026-10-08 01:27:58 +0200 |
|---|---|---|
| committer | Christophe Besson <cbesson@gmail.com> | 2026-10-08 01:27:58 +0200 |
| commit | c27e04c88557716bbe8e42b9174ba9b07f90facf (patch) | |
| tree | a1a56edbede2bba9bb0cc99b81b98570b8313915 /docs | |
| parent | cdd5fd52e981c4c59643e7dee705b6c55acae68e (diff) | |
| download | meshbay-0.19.tar.gz | |
perf(node): sample 9 MB with the size above 9 MB, keep known ids0.19
hash_version 3: size + first 4 MB + last 4 MB + 1 MB at the middle,
5.5x faster cold on a USB disk than the 45 MB sample. The cache now
serves a hit under whatever version it holds, so no existing id moves.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Diffstat (limited to 'docs')
| -rw-r--r-- | docs/MESHBAY_DESIGN.md | 33 | ||||
| -rw-r--r-- | docs/USERGUIDE.md | 2 | ||||
| -rw-r--r-- | docs/playlists.md | 2 |
3 files changed, 22 insertions, 15 deletions
diff --git a/docs/MESHBAY_DESIGN.md b/docs/MESHBAY_DESIGN.md index 23ced38..18b23cf 100644 --- a/docs/MESHBAY_DESIGN.md +++ b/docs/MESHBAY_DESIGN.md @@ -1620,27 +1620,34 @@ report ten files and index nine, and it decides how reconciliation must work: > sweep, forever — rewriting the entry, bumping the version and pushing an index > update to every connected peer. -**Hashing is partial above 40 MB.** A hash exists for content identity, and a 4 GB +**Hashing is partial above 9 MB.** A hash exists for content identity, and a 4 GB file does not need 4 GB of I/O to be identified with overwhelming probability: | Size | Method | `hash_version` | |---|---|---| -| ≤ 40 MB | full read | `1` | -| > 40 MB | blake3 over the first 20 MB ‖ last 20 MB ‖ 5 MB at the midpoint | `2` | +| ≤ 9 MB | full read | `1` | +| > 9 MB | blake3 over the size (8 bytes LE) ‖ first 4 MB ‖ last 4 MB ‖ 1 MB at the midpoint | `3` | +| (> 40 MB, up to 0.19.0) | first 20 MB ‖ last 20 MB ‖ 5 MB at the midpoint | `2`, no longer computed | Head and tail catch container headers, trailers and files that differ only at one -end; the mid-sample catches files sharing a header and trailer. Below the -threshold a partial read would sample the whole file anyway, so the full path is -simpler and produces the same value — which is what keeps small files -cross-comparable between nodes of different versions. A large file indexed by a -node of each version produces different ids and does not merge in cross-group -search; that resolves itself when both upgrade, and is the accepted cost of not -reading 4 TB to build a library. +end; the mid-sample catches files sharing a header and trailer; the size separates +files whose samples agree, such as two preallocated downloads still full of zeros. +No sample size catches an edit in the middle of a large file, so a larger sample +buys nothing in identity. On a spinning disk it buys only time: the cost is the +three seeks. Measured cold on a USB drive, 45 MB took 516 ms a file, 9 MB 94 ms +and 3 MB 79 ms. The threshold equals the sample, so no file costs more to read +than a sample would. + +**A file keeps the id it was first given.** The hash cache serves a hit under +whatever `hash_version` it was computed with, because TMDB matches and manual +corrections, thumbnails and members' playlists are keyed by the id. Only a file +the node has not hashed before, or one whose size or mtime changed, gets the +current scheme. The same file indexed by nodes of different versions can +therefore carry different ids and not merge in cross-group search: the accepted +cost of not orphaning anyone's references. `hash_version` is an additive index field with a default, so an entry written -before it deserialises correctly and needs no protocol bump. The hash cache -carries the column and auto-migrates on open; large files are re-hashed lazily on -the first scan after an upgrade. +before it deserialises correctly and needs no protocol bump. **Periodic reconciliation is mandatory on every platform**, not a backstop: `ReadDirectoryChangesW` drops events under load on Windows, and inotify is diff --git a/docs/USERGUIDE.md b/docs/USERGUIDE.md index adca9e3..2c37f56 100644 --- a/docs/USERGUIDE.md +++ b/docs/USERGUIDE.md @@ -549,7 +549,7 @@ The node watches its directories and re-checks them periodically — the periodi pass is not a backstop, it is required, because filesystem events are dropped under load and are unreliable on network and FUSE mounts. -Files over 40 MB are identified by reading their beginning, end and middle +Files over 9 MB are identified by their size and by reading their beginning, end and middle rather than the whole file. A 4 TB library does not need 4 TB of reading to be catalogued. diff --git a/docs/playlists.md b/docs/playlists.md index f1f6b67..05e6776 100644 --- a/docs/playlists.md +++ b/docs/playlists.md @@ -447,7 +447,7 @@ Each field prevents a specific failure: playlist is a column of hex strings. A display problem, solved by copying four small strings. - **`hv`** — the index already has two hashing schemes (`protocol.py:297`: - 1 = whole file, 2 = 45 MB sample). A re-hash would orphan every entry in every + 1 = whole file, 2 = 45 MB sample, 3 = size + 9 MB sample). A re-hash would orphan every entry in every playlist, silently and all at once. - **`p`** — content addressing survives a move; a path survives a re-encode. Keeping both means either can repair the other: on a sight of the live index, |