summaryrefslogtreecommitdiffstats
path: root/docs/indexing-v2.md
diff options
context:
space:
mode:
Diffstat (limited to 'docs/indexing-v2.md')
-rw-r--r--docs/indexing-v2.md324
1 files changed, 0 insertions, 324 deletions
diff --git a/docs/indexing-v2.md b/docs/indexing-v2.md
deleted file mode 100644
index a3522de..0000000
--- a/docs/indexing-v2.md
+++ /dev/null
@@ -1,324 +0,0 @@
-# Indexing v2 — Partial-read hashing for large files
-
-> **Superseded by `MESHBAY_DESIGN.md`.** This was the partial-read hashing; its design
-> content now lives in §6.3.
->
-> It is kept because code comments, tests and other documents cite its
-> sections and its labels, and because it records reasoning a synthesis
-> compresses. **Where it disagrees with `MESHBAY_DESIGN.md`, the design
-> document is right; where either disagrees with the code, the code is.**
-> `MESHBAY_DESIGN.md` §16 maps every section reference here onto its
-> replacement, and §13 defines every label.
-
-> Status: **built.** `hash_version` and partial-read hashing are in the indexer,
-> the protocol and the hash cache. Kept as the decision record and for the
-> migration notes; this header said "not built" long after it shipped.
-
----
-
-## 0. Problem
-
-The current indexer reads every file in full to compute its blake3 content hash (`id`).
-For a large media library (multi-terabyte, thousands of files), this means:
-
-- **Time**: initial indexing takes tens of minutes to hours.
-- **Disk I/O**: every byte of every file is read, which wears SSDs and saturates spinning
- drives for the entire duration. A USB hard drive serving a 4 TB library is pegged for
- over an hour.
-- **Blocking**: no connected peer receives a usable index until the full scan finishes.
-
-The hash exists for **content identity** (deduplication, cross-group search, file
-requests). A 4 GB film does not need 4 GB of I/O to be identified with overwhelming
-probability — 45 MB of well-chosen samples suffice.
-
----
-
-## 1. Design
-
-### 1.1 Hashing rules
-
-| File size | Method | `hash_version` |
-|--------------------|-----------------------------------------------------|-----------------|
-| <= 40 MB | Full read, blake3 of entire content (unchanged) | `1` |
-| > 40 MB | Partial read, blake3 of 45 MB sampled (see below) | `2` |
-
-**Partial-read algorithm (hash_version 2):**
-
-Given a file of `S` bytes where `S > 40 MB`:
-
-1. Read the first **20 MB** (bytes `[0, 20 MB)`).
-2. Append the last **20 MB** (bytes `[S - 20 MB, S)`).
-3. Append **5 MB** starting at **50% of the file** (bytes `[S // 2, S // 2 + 5 MB)`).
-4. Compute `blake3(concatenation of the three regions)`.
-
-The three regions may overlap for files just above 40 MB. This is fine — the concatenation
-is deterministic for a given file, which is the only property that matters.
-
-**Why these offsets.** Head and tail catch container headers, trailers, and the common case
-of files that differ only at one end (re-encoded, re-muxed, appended). The mid-sample
-catches files that share a header and trailer but differ in content (same container,
-different media stream).
-
-**Why 40 MB threshold.** Below 40 MB the partial read would sample the entire file anyway
-(head + tail >= file size), so the full-read path is both simpler and produces the same
-result. The boundary is inclusive: a 40 MB file is read in full.
-
-### 1.2 `hash_version` field
-
-A new field on `IndexEntry`:
-
-```
-hash_version: int = 1
-```
-
-- `1` — the `id` is blake3 of the full file content. This is the only value any existing
- node has ever produced.
-- `2` — the `id` is blake3 of the 45 MB partial sample described above.
-
-**For files <= 40 MB on a v2 node, `hash_version` stays `1`.** The hash is identical to
-what a v1 node produces, because both read the file in full. This preserves cross-group
-search compatibility for small files across v1 and v2 nodes.
-
-**For files > 40 MB on a v2 node, `hash_version` is `2`.** The hash is different from
-what a v1 node would produce for the same file. This is the accepted side effect.
-
-### 1.3 Backward compatibility
-
-| Scenario | Behaviour |
-|---|---|
-| v2 node sends `hash_version` to v1 client | Client ignores unknown field (JS objects are open) |
-| v1 node sends entries without `hash_version` | Client/consumer treats it as `1` (dataclass default) |
-| v2 `IndexEntry(**e)` where `e` lacks `hash_version` | Uses default `1` — existing serialized indexes deserialize correctly |
-| Cross-group search: same file, one node v1, one node v2 | Different `id` for files > 40 MB — not merged. Accepted |
-| Cross-group search: same small file, mixed nodes | Same `id` (both `hash_version=1`) — merged correctly |
-| Hub tables (`swarm_sources`, `content_blocklist`, `content_reports`) | Store `content_hash` as an opaque string. No change needed |
-| `GroupIndex.serialize()` / `deserialize()` | `asdict(e)` includes `hash_version`; `IndexEntry(**e)` with default handles missing field |
-
-**Nothing breaks.** A v1 node's data remains valid. A v2 node produces correct new hashes.
-Mixed v1/v2 environments work, with the documented search side effect.
-
----
-
-## 2. Affected components
-
-### 2.1 `meshbay_common` — `protocol.py`
-
-| Change | Detail |
-|---|---|
-| `IndexEntry` dataclass | Add `hash_version: int = 1` field |
-| `index_entry_wire()` | Add `"hash_version": e.hash_version` to the wire dict |
-
-### 2.2 `meshbay_node` — `indexer/indexer.py`
-
-| Change | Detail |
-|---|---|
-| Constants | `_PARTIAL_THRESHOLD = 40 * 1024 * 1024`, `_PARTIAL_HEAD = 20 * 1024 * 1024`, `_PARTIAL_TAIL = 20 * 1024 * 1024`, `_PARTIAL_MID = 5 * 1024 * 1024` |
-| `_scan_file()` | After stat, if `size > _PARTIAL_THRESHOLD`: use partial-read blake3. Set `hash_version=2` on the returned `IndexEntry`. Otherwise: unchanged (full read, `hash_version=1`) |
-| `_hash_or_cached()` | Compute expected `hash_version` from file size. Pass it to cache `lookup()`. Store it in cache `put()`. Set it on the returned `IndexEntry` |
-
-**`_scan_file` partial-read implementation:**
-
-```python
-def _partial_hash(file_path: Path, size: int) -> str:
- hasher = blake3.blake3()
- with open(long_path(file_path), "rb") as f:
- _feed(hasher, f, _PARTIAL_HEAD)
- f.seek(size - _PARTIAL_TAIL)
- _feed(hasher, f, _PARTIAL_TAIL)
- f.seek(size // 2)
- _feed(hasher, f, _PARTIAL_MID)
- return hasher.hexdigest()
-
-def _feed(hasher, f, nbytes: int) -> None:
- remaining = nbytes
- while remaining > 0:
- chunk = f.read(min(_HASH_CHUNK, remaining))
- if not chunk:
- break
- hasher.update(chunk)
- remaining -= len(chunk)
-```
-
-### 2.3 `meshbay_node` — `indexer/cache.py`
-
-| Change | Detail |
-|---|---|
-| Schema | Add `hash_version INTEGER NOT NULL DEFAULT 1` column to `files` table |
-| `_SCHEMA` | New databases get the column via `CREATE TABLE` |
-| `_MIGRATE_V2` | `ALTER TABLE files ADD COLUMN hash_version INTEGER NOT NULL DEFAULT 1` — applied on `open()` if the column does not exist |
-| `lookup()` | Add `hash_version` parameter. WHERE clause becomes `path = ? AND size = ? AND mtime = ? AND hash_version = ?` |
-| `put()` | Add `hash_version` parameter. INSERT includes `hash_version` |
-| `CachedEntry` | Add `hash_version: int` field |
-
-**Auto-migration on open:** `IndexCache.open()` runs the ALTER TABLE inside a try/except
-(column already exists → no-op). This way a node upgrade just works — no manual step
-needed for the cache.
-
-### 2.4 `meshbay_node` — `indexer/group_index.py`
-
-No code change needed. `serialize()` calls `asdict(e)` which includes `hash_version`.
-`deserialize()` calls `IndexEntry(**e)` which uses the default `1` for entries written
-before v2.
-
-### 2.5 `meshbay_node` — `transport/wire.py`
-
-No code change needed. `index_sync_message()` and `index_delta_message()` call
-`index_entry_wire()` which is updated in §2.1.
-
-### 2.6 Client side (browser / desktop)
-
-**No code changes required.** Entries are JavaScript objects; the extra `hash_version`
-field is carried through without needing explicit handling:
-
-- `transport.js` — `_applyIndexMessage()` passes the opened payload through. `hash_version`
- rides along on each entry object.
-- `group-page.js` — `applyIndex()` / `applyIndexDelta()` store entries as-is in state
- and in IndexedDB.
-- `hub-client.js` — IndexedDB cache stores entry objects verbatim.
-- `source-merge.js` — merges by `id`. Different hashes naturally don't merge.
-- `search-page.js` — aggregates entries from cached indexes. No change.
-- `files-app.js` — renders entries. `hash_version` is ignored.
-
-### 2.7 Hub side
-
-**No code changes required.** `swarm_sources.content_hash`, `content_reports.content_hash`,
-`content_blocklist.content_hash` are opaque `String(64)` columns. They store whatever
-blake3 hex the node provides. No hub migration needed.
-
-### 2.8 Tests
-
-| Test | What it verifies |
-|---|---|
-| `test_indexer.py` — new cases | Partial hash for file > 40 MB produces `hash_version=2`. File <= 40 MB produces `hash_version=1`. Partial hash is deterministic. Partial hash differs from full hash for the same large file |
-| `test_indexer.py` — existing cases | All existing tests still pass (small files, type detection, enrichment, etc.) |
-| `test_index_cache.py` — new cases | Cache lookup with `hash_version` match. Cache miss when `hash_version` differs. Schema migration from v1 cache |
-| `test_index_cache.py` — existing cases | Unchanged behaviour for v1 entries |
-| `test_indexer.py` — round-trip | `GroupIndex.serialize()` → `deserialize()` preserves `hash_version` on entries |
-| `test_index_seal_client.py` | Wire format includes `hash_version`, old entries without it deserialize as v1 |
-
----
-
-## 3. Migration
-
-### 3.1 What actually needs migrating
-
-The node has two relevant stores:
-
-| Store | Location | Content | Migration |
-|---|---|---|---|
-| `index_cache.db` | `data_dir/index_cache.db` | `(path, size, mtime) → hash` accelerator | Add `hash_version` column |
-| In-memory `GroupIndex` | rebuilt from disk on every startup | Current file listing | No migration — rebuilt on next scan |
-
-**The IndexCache is the only persistent store that needs a schema change.** The GroupIndex
-is rebuilt by scanning the filesystem on every daemon start. Once the code uses v2 hashing,
-the next startup produces v2 hashes for large files automatically.
-
-**The cache auto-migrates.** `IndexCache.open()` adds the `hash_version` column if missing.
-Existing rows get `DEFAULT 1`. When the v2 indexer looks up a large file with
-`hash_version=2`, the cached v1 entry won't match (different hash_version in WHERE), so
-the file is re-hashed with the partial algorithm and the new entry is written with
-`hash_version=2`.
-
-This means: **large files are re-hashed lazily on first scan after upgrade.** The first
-scan after upgrading to v2 re-reads 45 MB per large file instead of the full content —
-already much faster than v1's full read.
-
-### 3.2 Migration script — `QE/migration/migrate_index_v2.sh`
-
-A standalone bash script (not in git — QE/ is gitignored) that:
-
-1. Detects the node's `data_dir` from `node.toml` (default `~/.local/share/meshbay-node/`)
-2. Checks that `index_cache.db` exists
-3. Runs `ALTER TABLE files ADD COLUMN hash_version INTEGER NOT NULL DEFAULT 1`
-4. Optionally (`--purge-large`) deletes cache entries for files > 40 MB, forcing immediate
- re-hash on next scan instead of lazy migration
-5. Reports what it did
-
-**Works on Ubuntu and Fedora** — uses only `sqlite3` (present by default on both) and
-standard bash.
-
-**Not strictly required** if the code's auto-migration in `IndexCache.open()` is
-implemented. The script exists for operators who want to:
-- Verify the schema change before restarting
-- Force a clean re-hash of all large files in one pass
-- Run the migration on a machine where the node isn't installed yet (preparing a data dir)
-
-### 3.3 What the operator does
-
-1. Update the node package (or `pip install -e` in dev)
-2. (Optional) Run `QE/migration/migrate_index_v2.sh` to preview or force the migration
-3. Restart the node daemon
-4. The first scan re-hashes files > 40 MB with partial reads — much faster than before
-
-No client-side action needed. No hub-side action needed.
-
----
-
-## 4. What does NOT change
-
-- **Files <= 40 MB** — identical hashing, identical `id`, `hash_version=1`.
-- **Enrichment** (thumbnails, duration, metadata) — unchanged, still runs after hashing.
-- **File downloads** — `file_req` uses the current `id` from the index. After re-indexing,
- clients get the new index with new hashes and request accordingly.
-- **Watchdog / reconciliation** — filesystem events trigger the same code paths. New or
- modified files are hashed with the appropriate method based on size.
-- **Index encryption / signing** — `GroupIndex.serialize()` and the GEK-sealed wire
- messages are unchanged. `hash_version` rides inside each entry via `asdict()`.
-- **Index delta computation** — `GroupIndex.diff()` compares entries by all fields
- (dataclass `__eq__`). A re-indexed file whose hash changed (v1 → v2) appears as a
- deletion of the old id + addition of the new id, which is correct.
-- **Hub tables** — opaque content_hash storage, untouched.
-- **MNP version** — this is an additive field on index entries. No protocol version bump
- needed. An older node that doesn't send `hash_version` is handled by the default.
-
----
-
-## 5. Side effects — documented and accepted
-
-1. **Cross-group search**: a file > 40 MB indexed on a v1 node and a v2 node produces
- different `id`s. The search page will not merge them as the same file. The user sees
- two entries instead of one, each from its own group. This resolves itself when both
- nodes upgrade.
-
-2. **First scan after upgrade**: files > 40 MB are re-hashed. With v2 this reads 45 MB
- per file (not the full content), so the re-index is fast — but it is not instant. A
- 4 TB library with 1000 large files reads ~44 GB instead of 4 TB.
-
-3. **`content_hash` drift on hub tables**: if a public group's node upgrades, the hashes
- it registers in `swarm_sources` change for large files. Old entries with v1 hashes
- become stale. The swarm registration mechanism's `last_seen` update handles this — stale
- entries age out. No explicit cleanup needed.
-
-4. **Index version bump**: re-indexing sets `GroupIndex.version = int(time.time())`,
- which triggers a full `index_sync` to all connected peers. This is the normal path for
- any index change — it is not new load.
-
----
-
-## 6. Indexing trigger points — verified safe
-
-| Trigger | Location | Impact of v2 |
-|---|---|---|
-| Daemon startup — initial scan | `daemon.py:648-674` | Uses `_hash_or_cached()` which applies v2 rules. Safe |
-| Create Group wizard — Step 3 | `create-group-page.js:244-248` polls `index-status` | No change — polls progress, doesn't control hashing |
-| Settings — add root | `group-settings.js:909-927` | Triggers `retarget()` → scan → `_hash_or_cached()`. Safe |
-| Watchdog — file created/modified | `indexer.py:797-816` | Calls `_update_entry()` → `_hash_or_cached()`. Safe |
-| Reconciliation loop | `indexer.py:513-650` | Calls `_sweep_available_roots()` → `_hash_or_cached()`. Safe |
-| Hot reload — `_reload_config()` | `daemon.py:790-829` | Creates new `DirectoryIndexer` with v2 code. Safe |
-
----
-
-## 7. Implementation order
-
-1. **`protocol.py`** — add `hash_version` field to `IndexEntry` and `index_entry_wire()`
-2. **`cache.py`** — add `hash_version` to schema, auto-migrate on open, update lookup/put
-3. **`indexer.py`** — implement `_partial_hash()`, update `_scan_file()` and
- `_hash_or_cached()`
-4. **Tests** — new test cases for partial hashing, cache versioning, wire round-trip
-5. **`QE/migration/migrate_index_v2.sh`** — standalone migration script
-6. **Manual test** — run a node with a mixed library (small + large files), verify:
- - Small files: same hash as before, `hash_version=1`
- - Large files: different hash, `hash_version=2`, 45 MB read
- - Cross-group search: small files merge, large files don't (across v1/v2 nodes)
- - Cache hit on second scan: no re-read
- - Index sync to connected peers: entries carry `hash_version`