summaryrefslogtreecommitdiffstats
path: root/docs/MESHBAY_DESIGN.md
diff options
context:
space:
mode:
Diffstat (limited to 'docs/MESHBAY_DESIGN.md')
-rw-r--r--docs/MESHBAY_DESIGN.md33
1 files changed, 20 insertions, 13 deletions
diff --git a/docs/MESHBAY_DESIGN.md b/docs/MESHBAY_DESIGN.md
index 23ced38..18b23cf 100644
--- a/docs/MESHBAY_DESIGN.md
+++ b/docs/MESHBAY_DESIGN.md
@@ -1620,27 +1620,34 @@ report ten files and index nine, and it decides how reconciliation must work:
> sweep, forever — rewriting the entry, bumping the version and pushing an index
> update to every connected peer.
-**Hashing is partial above 40 MB.** A hash exists for content identity, and a 4 GB
+**Hashing is partial above 9 MB.** A hash exists for content identity, and a 4 GB
file does not need 4 GB of I/O to be identified with overwhelming probability:
| Size | Method | `hash_version` |
|---|---|---|
-| ≤ 40 MB | full read | `1` |
-| > 40 MB | blake3 over the first 20 MB ‖ last 20 MB ‖ 5 MB at the midpoint | `2` |
+| ≤ 9 MB | full read | `1` |
+| > 9 MB | blake3 over the size (8 bytes LE) ‖ first 4 MB ‖ last 4 MB ‖ 1 MB at the midpoint | `3` |
+| (> 40 MB, up to 0.19.0) | first 20 MB ‖ last 20 MB ‖ 5 MB at the midpoint | `2`, no longer computed |
Head and tail catch container headers, trailers and files that differ only at one
-end; the mid-sample catches files sharing a header and trailer. Below the
-threshold a partial read would sample the whole file anyway, so the full path is
-simpler and produces the same value — which is what keeps small files
-cross-comparable between nodes of different versions. A large file indexed by a
-node of each version produces different ids and does not merge in cross-group
-search; that resolves itself when both upgrade, and is the accepted cost of not
-reading 4 TB to build a library.
+end; the mid-sample catches files sharing a header and trailer; the size separates
+files whose samples agree, such as two preallocated downloads still full of zeros.
+No sample size catches an edit in the middle of a large file, so a larger sample
+buys nothing in identity. On a spinning disk it buys only time: the cost is the
+three seeks. Measured cold on a USB drive, 45 MB took 516 ms a file, 9 MB 94 ms
+and 3 MB 79 ms. The threshold equals the sample, so no file costs more to read
+than a sample would.
+
+**A file keeps the id it was first given.** The hash cache serves a hit under
+whatever `hash_version` it was computed with, because TMDB matches and manual
+corrections, thumbnails and members' playlists are keyed by the id. Only a file
+the node has not hashed before, or one whose size or mtime changed, gets the
+current scheme. The same file indexed by nodes of different versions can
+therefore carry different ids and not merge in cross-group search: the accepted
+cost of not orphaning anyone's references.
`hash_version` is an additive index field with a default, so an entry written
-before it deserialises correctly and needs no protocol bump. The hash cache
-carries the column and auto-migrates on open; large files are re-hashed lazily on
-the first scan after an upgrade.
+before it deserialises correctly and needs no protocol bump.
**Periodic reconciliation is mandatory on every platform**, not a backstop:
`ReadDirectoryChangesW` drops events under load on Windows, and inotify is