summaryrefslogtreecommitdiffstats
path: root/docs/indexing-v2.md
diff options
context:
space:
mode:
authorChristophe Besson <cbesson@gmail.com>2026-09-11 00:19:06 +0200
committerChristophe Besson <cbesson@gmail.com>2026-09-11 00:19:06 +0200
commitf059cb118c556d1f0279350507f74b8a47d5a98a (patch)
tree9a97762a844038a06134b4b7dcead1758477dfc1 /docs/indexing-v2.md
parentb045ba0010d69360b6a0265eb7c73a07900fe328 (diff)
downloadmeshbay-f059cb118c556d1f0279350507f74b8a47d5a98a.tar.gz
docs: remove the documents MESHBAY_DESIGN.md replaces
Twenty-four files, about 17 000 lines: the two architecture drafts, the three security reviews, eleven design notes, the roadmap, the decisions file, the v1–v4 archive, the deprecated user guide and the stale quickstart. Their content is in MESHBAY_DESIGN.md, and git history holds the originals. The reason to delete rather than keep bannered: a document that is superseded but present still gets read, and a reader cannot always tell which of two accounts of one mechanism is the live one. That was the argument for retiring the user guide rather than repairing it, and it applies to the whole set. What made this safe is the concordance. Roughly 290 comments and docstrings cite these files by section — `musicbay.md §6`, `mediacenter.md §5.5`, `draft-v6 §2.11` — and section 16 maps every one onto its replacement, so not a single comment needs editing to stay followable. It now says plainly that the files are gone and where to recover them, and it gained rows for the three reviews (their findings are section 13), and for the two guides. Four kept documents pointed into the set and were repointed first: `playlists.md` (nine references — it is a live proposal and must not dangle), `WINDOWS-PORT.md`, and CLAUDE.md's example. No dangling reference remains outside section 16. Two files were dropped from the list after checking what they hold. `HTTPS.md` is an operational runbook — Caddy, certificate renewal, DNS, troubleshooting — and MESHBAY_DESIGN.md deliberately covers no operations, so nothing would replace it; the versioned Caddyfile is the config, not the procedure. `cast-smart-tv.md` is the plan for the unbuilt DLNA phase of a feature whose first two phases ship, and section 11.4 summarises it in four lines rather than carrying the SSDP/UPnP work. There is no user guide now, and section 0.1 says so rather than leaving a reader to discover it. Suites green: 2258 passed, 4 skipped. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YVoHVCcfBqud6ZjG4db3y7
Diffstat (limited to 'docs/indexing-v2.md')
-rw-r--r--docs/indexing-v2.md324
1 files changed, 0 insertions, 324 deletions
diff --git a/docs/indexing-v2.md b/docs/indexing-v2.md
deleted file mode 100644
index a3522de..0000000
--- a/docs/indexing-v2.md
+++ /dev/null
@@ -1,324 +0,0 @@
-# Indexing v2 — Partial-read hashing for large files
-
-> **Superseded by `MESHBAY_DESIGN.md`.** This was the partial-read hashing; its design
-> content now lives in §6.3.
->
-> It is kept because code comments, tests and other documents cite its
-> sections and its labels, and because it records reasoning a synthesis
-> compresses. **Where it disagrees with `MESHBAY_DESIGN.md`, the design
-> document is right; where either disagrees with the code, the code is.**
-> `MESHBAY_DESIGN.md` §16 maps every section reference here onto its
-> replacement, and §13 defines every label.
-
-> Status: **built.** `hash_version` and partial-read hashing are in the indexer,
-> the protocol and the hash cache. Kept as the decision record and for the
-> migration notes; this header said "not built" long after it shipped.
-
----
-
-## 0. Problem
-
-The current indexer reads every file in full to compute its blake3 content hash (`id`).
-For a large media library (multi-terabyte, thousands of files), this means:
-
-- **Time**: initial indexing takes tens of minutes to hours.
-- **Disk I/O**: every byte of every file is read, which wears SSDs and saturates spinning
- drives for the entire duration. A USB hard drive serving a 4 TB library is pegged for
- over an hour.
-- **Blocking**: no connected peer receives a usable index until the full scan finishes.
-
-The hash exists for **content identity** (deduplication, cross-group search, file
-requests). A 4 GB film does not need 4 GB of I/O to be identified with overwhelming
-probability — 45 MB of well-chosen samples suffice.
-
----
-
-## 1. Design
-
-### 1.1 Hashing rules
-
-| File size | Method | `hash_version` |
-|--------------------|-----------------------------------------------------|-----------------|
-| <= 40 MB | Full read, blake3 of entire content (unchanged) | `1` |
-| > 40 MB | Partial read, blake3 of 45 MB sampled (see below) | `2` |
-
-**Partial-read algorithm (hash_version 2):**
-
-Given a file of `S` bytes where `S > 40 MB`:
-
-1. Read the first **20 MB** (bytes `[0, 20 MB)`).
-2. Append the last **20 MB** (bytes `[S - 20 MB, S)`).
-3. Append **5 MB** starting at **50% of the file** (bytes `[S // 2, S // 2 + 5 MB)`).
-4. Compute `blake3(concatenation of the three regions)`.
-
-The three regions may overlap for files just above 40 MB. This is fine — the concatenation
-is deterministic for a given file, which is the only property that matters.
-
-**Why these offsets.** Head and tail catch container headers, trailers, and the common case
-of files that differ only at one end (re-encoded, re-muxed, appended). The mid-sample
-catches files that share a header and trailer but differ in content (same container,
-different media stream).
-
-**Why 40 MB threshold.** Below 40 MB the partial read would sample the entire file anyway
-(head + tail >= file size), so the full-read path is both simpler and produces the same
-result. The boundary is inclusive: a 40 MB file is read in full.
-
-### 1.2 `hash_version` field
-
-A new field on `IndexEntry`:
-
-```
-hash_version: int = 1
-```
-
-- `1` — the `id` is blake3 of the full file content. This is the only value any existing
- node has ever produced.
-- `2` — the `id` is blake3 of the 45 MB partial sample described above.
-
-**For files <= 40 MB on a v2 node, `hash_version` stays `1`.** The hash is identical to
-what a v1 node produces, because both read the file in full. This preserves cross-group
-search compatibility for small files across v1 and v2 nodes.
-
-**For files > 40 MB on a v2 node, `hash_version` is `2`.** The hash is different from
-what a v1 node would produce for the same file. This is the accepted side effect.
-
-### 1.3 Backward compatibility
-
-| Scenario | Behaviour |
-|---|---|
-| v2 node sends `hash_version` to v1 client | Client ignores unknown field (JS objects are open) |
-| v1 node sends entries without `hash_version` | Client/consumer treats it as `1` (dataclass default) |
-| v2 `IndexEntry(**e)` where `e` lacks `hash_version` | Uses default `1` — existing serialized indexes deserialize correctly |
-| Cross-group search: same file, one node v1, one node v2 | Different `id` for files > 40 MB — not merged. Accepted |
-| Cross-group search: same small file, mixed nodes | Same `id` (both `hash_version=1`) — merged correctly |
-| Hub tables (`swarm_sources`, `content_blocklist`, `content_reports`) | Store `content_hash` as an opaque string. No change needed |
-| `GroupIndex.serialize()` / `deserialize()` | `asdict(e)` includes `hash_version`; `IndexEntry(**e)` with default handles missing field |
-
-**Nothing breaks.** A v1 node's data remains valid. A v2 node produces correct new hashes.
-Mixed v1/v2 environments work, with the documented search side effect.
-
----
-
-## 2. Affected components
-
-### 2.1 `meshbay_common` — `protocol.py`
-
-| Change | Detail |
-|---|---|
-| `IndexEntry` dataclass | Add `hash_version: int = 1` field |
-| `index_entry_wire()` | Add `"hash_version": e.hash_version` to the wire dict |
-
-### 2.2 `meshbay_node` — `indexer/indexer.py`
-
-| Change | Detail |
-|---|---|
-| Constants | `_PARTIAL_THRESHOLD = 40 * 1024 * 1024`, `_PARTIAL_HEAD = 20 * 1024 * 1024`, `_PARTIAL_TAIL = 20 * 1024 * 1024`, `_PARTIAL_MID = 5 * 1024 * 1024` |
-| `_scan_file()` | After stat, if `size > _PARTIAL_THRESHOLD`: use partial-read blake3. Set `hash_version=2` on the returned `IndexEntry`. Otherwise: unchanged (full read, `hash_version=1`) |
-| `_hash_or_cached()` | Compute expected `hash_version` from file size. Pass it to cache `lookup()`. Store it in cache `put()`. Set it on the returned `IndexEntry` |
-
-**`_scan_file` partial-read implementation:**
-
-```python
-def _partial_hash(file_path: Path, size: int) -> str:
- hasher = blake3.blake3()
- with open(long_path(file_path), "rb") as f:
- _feed(hasher, f, _PARTIAL_HEAD)
- f.seek(size - _PARTIAL_TAIL)
- _feed(hasher, f, _PARTIAL_TAIL)
- f.seek(size // 2)
- _feed(hasher, f, _PARTIAL_MID)
- return hasher.hexdigest()
-
-def _feed(hasher, f, nbytes: int) -> None:
- remaining = nbytes
- while remaining > 0:
- chunk = f.read(min(_HASH_CHUNK, remaining))
- if not chunk:
- break
- hasher.update(chunk)
- remaining -= len(chunk)
-```
-
-### 2.3 `meshbay_node` — `indexer/cache.py`
-
-| Change | Detail |
-|---|---|
-| Schema | Add `hash_version INTEGER NOT NULL DEFAULT 1` column to `files` table |
-| `_SCHEMA` | New databases get the column via `CREATE TABLE` |
-| `_MIGRATE_V2` | `ALTER TABLE files ADD COLUMN hash_version INTEGER NOT NULL DEFAULT 1` — applied on `open()` if the column does not exist |
-| `lookup()` | Add `hash_version` parameter. WHERE clause becomes `path = ? AND size = ? AND mtime = ? AND hash_version = ?` |
-| `put()` | Add `hash_version` parameter. INSERT includes `hash_version` |
-| `CachedEntry` | Add `hash_version: int` field |
-
-**Auto-migration on open:** `IndexCache.open()` runs the ALTER TABLE inside a try/except
-(column already exists → no-op). This way a node upgrade just works — no manual step
-needed for the cache.
-
-### 2.4 `meshbay_node` — `indexer/group_index.py`
-
-No code change needed. `serialize()` calls `asdict(e)` which includes `hash_version`.
-`deserialize()` calls `IndexEntry(**e)` which uses the default `1` for entries written
-before v2.
-
-### 2.5 `meshbay_node` — `transport/wire.py`
-
-No code change needed. `index_sync_message()` and `index_delta_message()` call
-`index_entry_wire()` which is updated in §2.1.
-
-### 2.6 Client side (browser / desktop)
-
-**No code changes required.** Entries are JavaScript objects; the extra `hash_version`
-field is carried through without needing explicit handling:
-
-- `transport.js` — `_applyIndexMessage()` passes the opened payload through. `hash_version`
- rides along on each entry object.
-- `group-page.js` — `applyIndex()` / `applyIndexDelta()` store entries as-is in state
- and in IndexedDB.
-- `hub-client.js` — IndexedDB cache stores entry objects verbatim.
-- `source-merge.js` — merges by `id`. Different hashes naturally don't merge.
-- `search-page.js` — aggregates entries from cached indexes. No change.
-- `files-app.js` — renders entries. `hash_version` is ignored.
-
-### 2.7 Hub side
-
-**No code changes required.** `swarm_sources.content_hash`, `content_reports.content_hash`,
-`content_blocklist.content_hash` are opaque `String(64)` columns. They store whatever
-blake3 hex the node provides. No hub migration needed.
-
-### 2.8 Tests
-
-| Test | What it verifies |
-|---|---|
-| `test_indexer.py` — new cases | Partial hash for file > 40 MB produces `hash_version=2`. File <= 40 MB produces `hash_version=1`. Partial hash is deterministic. Partial hash differs from full hash for the same large file |
-| `test_indexer.py` — existing cases | All existing tests still pass (small files, type detection, enrichment, etc.) |
-| `test_index_cache.py` — new cases | Cache lookup with `hash_version` match. Cache miss when `hash_version` differs. Schema migration from v1 cache |
-| `test_index_cache.py` — existing cases | Unchanged behaviour for v1 entries |
-| `test_indexer.py` — round-trip | `GroupIndex.serialize()` → `deserialize()` preserves `hash_version` on entries |
-| `test_index_seal_client.py` | Wire format includes `hash_version`, old entries without it deserialize as v1 |
-
----
-
-## 3. Migration
-
-### 3.1 What actually needs migrating
-
-The node has two relevant stores:
-
-| Store | Location | Content | Migration |
-|---|---|---|---|
-| `index_cache.db` | `data_dir/index_cache.db` | `(path, size, mtime) → hash` accelerator | Add `hash_version` column |
-| In-memory `GroupIndex` | rebuilt from disk on every startup | Current file listing | No migration — rebuilt on next scan |
-
-**The IndexCache is the only persistent store that needs a schema change.** The GroupIndex
-is rebuilt by scanning the filesystem on every daemon start. Once the code uses v2 hashing,
-the next startup produces v2 hashes for large files automatically.
-
-**The cache auto-migrates.** `IndexCache.open()` adds the `hash_version` column if missing.
-Existing rows get `DEFAULT 1`. When the v2 indexer looks up a large file with
-`hash_version=2`, the cached v1 entry won't match (different hash_version in WHERE), so
-the file is re-hashed with the partial algorithm and the new entry is written with
-`hash_version=2`.
-
-This means: **large files are re-hashed lazily on first scan after upgrade.** The first
-scan after upgrading to v2 re-reads 45 MB per large file instead of the full content —
-already much faster than v1's full read.
-
-### 3.2 Migration script — `QE/migration/migrate_index_v2.sh`
-
-A standalone bash script (not in git — QE/ is gitignored) that:
-
-1. Detects the node's `data_dir` from `node.toml` (default `~/.local/share/meshbay-node/`)
-2. Checks that `index_cache.db` exists
-3. Runs `ALTER TABLE files ADD COLUMN hash_version INTEGER NOT NULL DEFAULT 1`
-4. Optionally (`--purge-large`) deletes cache entries for files > 40 MB, forcing immediate
- re-hash on next scan instead of lazy migration
-5. Reports what it did
-
-**Works on Ubuntu and Fedora** — uses only `sqlite3` (present by default on both) and
-standard bash.
-
-**Not strictly required** if the code's auto-migration in `IndexCache.open()` is
-implemented. The script exists for operators who want to:
-- Verify the schema change before restarting
-- Force a clean re-hash of all large files in one pass
-- Run the migration on a machine where the node isn't installed yet (preparing a data dir)
-
-### 3.3 What the operator does
-
-1. Update the node package (or `pip install -e` in dev)
-2. (Optional) Run `QE/migration/migrate_index_v2.sh` to preview or force the migration
-3. Restart the node daemon
-4. The first scan re-hashes files > 40 MB with partial reads — much faster than before
-
-No client-side action needed. No hub-side action needed.
-
----
-
-## 4. What does NOT change
-
-- **Files <= 40 MB** — identical hashing, identical `id`, `hash_version=1`.
-- **Enrichment** (thumbnails, duration, metadata) — unchanged, still runs after hashing.
-- **File downloads** — `file_req` uses the current `id` from the index. After re-indexing,
- clients get the new index with new hashes and request accordingly.
-- **Watchdog / reconciliation** — filesystem events trigger the same code paths. New or
- modified files are hashed with the appropriate method based on size.
-- **Index encryption / signing** — `GroupIndex.serialize()` and the GEK-sealed wire
- messages are unchanged. `hash_version` rides inside each entry via `asdict()`.
-- **Index delta computation** — `GroupIndex.diff()` compares entries by all fields
- (dataclass `__eq__`). A re-indexed file whose hash changed (v1 → v2) appears as a
- deletion of the old id + addition of the new id, which is correct.
-- **Hub tables** — opaque content_hash storage, untouched.
-- **MNP version** — this is an additive field on index entries. No protocol version bump
- needed. An older node that doesn't send `hash_version` is handled by the default.
-
----
-
-## 5. Side effects — documented and accepted
-
-1. **Cross-group search**: a file > 40 MB indexed on a v1 node and a v2 node produces
- different `id`s. The search page will not merge them as the same file. The user sees
- two entries instead of one, each from its own group. This resolves itself when both
- nodes upgrade.
-
-2. **First scan after upgrade**: files > 40 MB are re-hashed. With v2 this reads 45 MB
- per file (not the full content), so the re-index is fast — but it is not instant. A
- 4 TB library with 1000 large files reads ~44 GB instead of 4 TB.
-
-3. **`content_hash` drift on hub tables**: if a public group's node upgrades, the hashes
- it registers in `swarm_sources` change for large files. Old entries with v1 hashes
- become stale. The swarm registration mechanism's `last_seen` update handles this — stale
- entries age out. No explicit cleanup needed.
-
-4. **Index version bump**: re-indexing sets `GroupIndex.version = int(time.time())`,
- which triggers a full `index_sync` to all connected peers. This is the normal path for
- any index change — it is not new load.
-
----
-
-## 6. Indexing trigger points — verified safe
-
-| Trigger | Location | Impact of v2 |
-|---|---|---|
-| Daemon startup — initial scan | `daemon.py:648-674` | Uses `_hash_or_cached()` which applies v2 rules. Safe |
-| Create Group wizard — Step 3 | `create-group-page.js:244-248` polls `index-status` | No change — polls progress, doesn't control hashing |
-| Settings — add root | `group-settings.js:909-927` | Triggers `retarget()` → scan → `_hash_or_cached()`. Safe |
-| Watchdog — file created/modified | `indexer.py:797-816` | Calls `_update_entry()` → `_hash_or_cached()`. Safe |
-| Reconciliation loop | `indexer.py:513-650` | Calls `_sweep_available_roots()` → `_hash_or_cached()`. Safe |
-| Hot reload — `_reload_config()` | `daemon.py:790-829` | Creates new `DirectoryIndexer` with v2 code. Safe |
-
----
-
-## 7. Implementation order
-
-1. **`protocol.py`** — add `hash_version` field to `IndexEntry` and `index_entry_wire()`
-2. **`cache.py`** — add `hash_version` to schema, auto-migrate on open, update lookup/put
-3. **`indexer.py`** — implement `_partial_hash()`, update `_scan_file()` and
- `_hash_or_cached()`
-4. **Tests** — new test cases for partial hashing, cache versioning, wire round-trip
-5. **`QE/migration/migrate_index_v2.sh`** — standalone migration script
-6. **Manual test** — run a node with a mixed library (small + large files), verify:
- - Small files: same hash as before, `hash_version=1`
- - Large files: different hash, `hash_version=2`, 45 MB read
- - Cross-group search: small files merge, large files don't (across v1/v2 nodes)
- - Cache hit on second scan: no re-read
- - Index sync to connected peers: entries carry `hash_version`