aboutsummaryrefslogtreecommitdiffstats
path: root/docs
diff options
context:
space:
mode:
Diffstat (limited to 'docs')
-rw-r--r--docs/MESHBAY_DESIGN.md33
-rw-r--r--docs/USERGUIDE.md2
-rw-r--r--docs/playlists.md2
3 files changed, 22 insertions, 15 deletions
diff --git a/docs/MESHBAY_DESIGN.md b/docs/MESHBAY_DESIGN.md
index 23ced38..18b23cf 100644
--- a/docs/MESHBAY_DESIGN.md
+++ b/docs/MESHBAY_DESIGN.md
@@ -1620,27 +1620,34 @@ report ten files and index nine, and it decides how reconciliation must work:
> sweep, forever — rewriting the entry, bumping the version and pushing an index
> update to every connected peer.
-**Hashing is partial above 40 MB.** A hash exists for content identity, and a 4 GB
+**Hashing is partial above 9 MB.** A hash exists for content identity, and a 4 GB
file does not need 4 GB of I/O to be identified with overwhelming probability:
| Size | Method | `hash_version` |
|---|---|---|
-| ≤ 40 MB | full read | `1` |
-| > 40 MB | blake3 over the first 20 MB ‖ last 20 MB ‖ 5 MB at the midpoint | `2` |
+| ≤ 9 MB | full read | `1` |
+| > 9 MB | blake3 over the size (8 bytes LE) ‖ first 4 MB ‖ last 4 MB ‖ 1 MB at the midpoint | `3` |
+| (> 40 MB, up to 0.19.0) | first 20 MB ‖ last 20 MB ‖ 5 MB at the midpoint | `2`, no longer computed |
Head and tail catch container headers, trailers and files that differ only at one
-end; the mid-sample catches files sharing a header and trailer. Below the
-threshold a partial read would sample the whole file anyway, so the full path is
-simpler and produces the same value — which is what keeps small files
-cross-comparable between nodes of different versions. A large file indexed by a
-node of each version produces different ids and does not merge in cross-group
-search; that resolves itself when both upgrade, and is the accepted cost of not
-reading 4 TB to build a library.
+end; the mid-sample catches files sharing a header and trailer; the size separates
+files whose samples agree, such as two preallocated downloads still full of zeros.
+No sample size catches an edit in the middle of a large file, so a larger sample
+buys nothing in identity. On a spinning disk it buys only time: the cost is the
+three seeks. Measured cold on a USB drive, 45 MB took 516 ms a file, 9 MB 94 ms
+and 3 MB 79 ms. The threshold equals the sample, so no file costs more to read
+than a sample would.
+
+**A file keeps the id it was first given.** The hash cache serves a hit under
+whatever `hash_version` it was computed with, because TMDB matches and manual
+corrections, thumbnails and members' playlists are keyed by the id. Only a file
+the node has not hashed before, or one whose size or mtime changed, gets the
+current scheme. The same file indexed by nodes of different versions can
+therefore carry different ids and not merge in cross-group search: the accepted
+cost of not orphaning anyone's references.
`hash_version` is an additive index field with a default, so an entry written
-before it deserialises correctly and needs no protocol bump. The hash cache
-carries the column and auto-migrates on open; large files are re-hashed lazily on
-the first scan after an upgrade.
+before it deserialises correctly and needs no protocol bump.
**Periodic reconciliation is mandatory on every platform**, not a backstop:
`ReadDirectoryChangesW` drops events under load on Windows, and inotify is
diff --git a/docs/USERGUIDE.md b/docs/USERGUIDE.md
index adca9e3..2c37f56 100644
--- a/docs/USERGUIDE.md
+++ b/docs/USERGUIDE.md
@@ -549,7 +549,7 @@ The node watches its directories and re-checks them periodically — the periodi
pass is not a backstop, it is required, because filesystem events are dropped
under load and are unreliable on network and FUSE mounts.
-Files over 40 MB are identified by reading their beginning, end and middle
+Files over 9 MB are identified by their size and by reading their beginning, end and middle
rather than the whole file. A 4 TB library does not need 4 TB of reading to be
catalogued.
diff --git a/docs/playlists.md b/docs/playlists.md
index f1f6b67..05e6776 100644
--- a/docs/playlists.md
+++ b/docs/playlists.md
@@ -447,7 +447,7 @@ Each field prevents a specific failure:
playlist is a column of hex strings. A display problem, solved by copying four
small strings.
- **`hv`** — the index already has two hashing schemes (`protocol.py:297`:
- 1 = whole file, 2 = 45 MB sample). A re-hash would orphan every entry in every
+ 1 = whole file, 2 = 45 MB sample, 3 = size + 9 MB sample). A re-hash would orphan every entry in every
playlist, silently and all at once.
- **`p`** — content addressing survives a move; a path survives a re-encode.
Keeping both means either can repair the other: on a sight of the live index,