diff options
Diffstat (limited to 'docs/MESHBAY_DESIGN.md')
| -rw-r--r-- | docs/MESHBAY_DESIGN.md | 33 |
1 files changed, 20 insertions, 13 deletions
diff --git a/docs/MESHBAY_DESIGN.md b/docs/MESHBAY_DESIGN.md index 23ced38..18b23cf 100644 --- a/docs/MESHBAY_DESIGN.md +++ b/docs/MESHBAY_DESIGN.md @@ -1620,27 +1620,34 @@ report ten files and index nine, and it decides how reconciliation must work: > sweep, forever — rewriting the entry, bumping the version and pushing an index > update to every connected peer. -**Hashing is partial above 40 MB.** A hash exists for content identity, and a 4 GB +**Hashing is partial above 9 MB.** A hash exists for content identity, and a 4 GB file does not need 4 GB of I/O to be identified with overwhelming probability: | Size | Method | `hash_version` | |---|---|---| -| ≤ 40 MB | full read | `1` | -| > 40 MB | blake3 over the first 20 MB ‖ last 20 MB ‖ 5 MB at the midpoint | `2` | +| ≤ 9 MB | full read | `1` | +| > 9 MB | blake3 over the size (8 bytes LE) ‖ first 4 MB ‖ last 4 MB ‖ 1 MB at the midpoint | `3` | +| (> 40 MB, up to 0.19.0) | first 20 MB ‖ last 20 MB ‖ 5 MB at the midpoint | `2`, no longer computed | Head and tail catch container headers, trailers and files that differ only at one -end; the mid-sample catches files sharing a header and trailer. Below the -threshold a partial read would sample the whole file anyway, so the full path is -simpler and produces the same value — which is what keeps small files -cross-comparable between nodes of different versions. A large file indexed by a -node of each version produces different ids and does not merge in cross-group -search; that resolves itself when both upgrade, and is the accepted cost of not -reading 4 TB to build a library. +end; the mid-sample catches files sharing a header and trailer; the size separates +files whose samples agree, such as two preallocated downloads still full of zeros. +No sample size catches an edit in the middle of a large file, so a larger sample +buys nothing in identity. On a spinning disk it buys only time: the cost is the +three seeks. Measured cold on a USB drive, 45 MB took 516 ms a file, 9 MB 94 ms +and 3 MB 79 ms. The threshold equals the sample, so no file costs more to read +than a sample would. + +**A file keeps the id it was first given.** The hash cache serves a hit under +whatever `hash_version` it was computed with, because TMDB matches and manual +corrections, thumbnails and members' playlists are keyed by the id. Only a file +the node has not hashed before, or one whose size or mtime changed, gets the +current scheme. The same file indexed by nodes of different versions can +therefore carry different ids and not merge in cross-group search: the accepted +cost of not orphaning anyone's references. `hash_version` is an additive index field with a default, so an entry written -before it deserialises correctly and needs no protocol bump. The hash cache -carries the column and auto-migrates on open; large files are re-hashed lazily on -the first scan after an upgrade. +before it deserialises correctly and needs no protocol bump. **Periodic reconciliation is mandatory on every platform**, not a backstop: `ReadDirectoryChangesW` drops events under load on Windows, and inotify is |