summaryrefslogtreecommitdiffstats
path: root/docs/indexing-v2.md
blob: a3522de3b7f3e89edc826f7ca1ac32970e17681d (plain) (blame)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
# Indexing v2 — Partial-read hashing for large files

> **Superseded by `MESHBAY_DESIGN.md`.** This was the partial-read hashing; its design
> content now lives in §6.3.
>
> It is kept because code comments, tests and other documents cite its
> sections and its labels, and because it records reasoning a synthesis
> compresses. **Where it disagrees with `MESHBAY_DESIGN.md`, the design
> document is right; where either disagrees with the code, the code is.**
> `MESHBAY_DESIGN.md` §16 maps every section reference here onto its
> replacement, and §13 defines every label.

> Status: **built.** `hash_version` and partial-read hashing are in the indexer,
> the protocol and the hash cache. Kept as the decision record and for the
> migration notes; this header said "not built" long after it shipped.

---

## 0. Problem

The current indexer reads every file in full to compute its blake3 content hash (`id`).
For a large media library (multi-terabyte, thousands of files), this means:

- **Time**: initial indexing takes tens of minutes to hours.
- **Disk I/O**: every byte of every file is read, which wears SSDs and saturates spinning
  drives for the entire duration. A USB hard drive serving a 4 TB library is pegged for
  over an hour.
- **Blocking**: no connected peer receives a usable index until the full scan finishes.

The hash exists for **content identity** (deduplication, cross-group search, file
requests). A 4 GB film does not need 4 GB of I/O to be identified with overwhelming
probability — 45 MB of well-chosen samples suffice.

---

## 1. Design

### 1.1 Hashing rules

| File size          | Method                                              | `hash_version` |
|--------------------|-----------------------------------------------------|-----------------|
| <= 40 MB           | Full read, blake3 of entire content (unchanged)     | `1`             |
| > 40 MB            | Partial read, blake3 of 45 MB sampled (see below)   | `2`             |

**Partial-read algorithm (hash_version 2):**

Given a file of `S` bytes where `S > 40 MB`:

1. Read the first **20 MB** (bytes `[0, 20 MB)`).
2. Append the last **20 MB** (bytes `[S - 20 MB, S)`).
3. Append **5 MB** starting at **50% of the file** (bytes `[S // 2, S // 2 + 5 MB)`).
4. Compute `blake3(concatenation of the three regions)`.

The three regions may overlap for files just above 40 MB. This is fine — the concatenation
is deterministic for a given file, which is the only property that matters.

**Why these offsets.** Head and tail catch container headers, trailers, and the common case
of files that differ only at one end (re-encoded, re-muxed, appended). The mid-sample
catches files that share a header and trailer but differ in content (same container,
different media stream).

**Why 40 MB threshold.** Below 40 MB the partial read would sample the entire file anyway
(head + tail >= file size), so the full-read path is both simpler and produces the same
result. The boundary is inclusive: a 40 MB file is read in full.

### 1.2 `hash_version` field

A new field on `IndexEntry`:

```
hash_version: int = 1
```

- `1` — the `id` is blake3 of the full file content. This is the only value any existing
  node has ever produced.
- `2` — the `id` is blake3 of the 45 MB partial sample described above.

**For files <= 40 MB on a v2 node, `hash_version` stays `1`.** The hash is identical to
what a v1 node produces, because both read the file in full. This preserves cross-group
search compatibility for small files across v1 and v2 nodes.

**For files > 40 MB on a v2 node, `hash_version` is `2`.** The hash is different from
what a v1 node would produce for the same file. This is the accepted side effect.

### 1.3 Backward compatibility

| Scenario | Behaviour |
|---|---|
| v2 node sends `hash_version` to v1 client | Client ignores unknown field (JS objects are open) |
| v1 node sends entries without `hash_version` | Client/consumer treats it as `1` (dataclass default) |
| v2 `IndexEntry(**e)` where `e` lacks `hash_version` | Uses default `1` — existing serialized indexes deserialize correctly |
| Cross-group search: same file, one node v1, one node v2 | Different `id` for files > 40 MB — not merged. Accepted |
| Cross-group search: same small file, mixed nodes | Same `id` (both `hash_version=1`) — merged correctly |
| Hub tables (`swarm_sources`, `content_blocklist`, `content_reports`) | Store `content_hash` as an opaque string. No change needed |
| `GroupIndex.serialize()` / `deserialize()` | `asdict(e)` includes `hash_version`; `IndexEntry(**e)` with default handles missing field |

**Nothing breaks.** A v1 node's data remains valid. A v2 node produces correct new hashes.
Mixed v1/v2 environments work, with the documented search side effect.

---

## 2. Affected components

### 2.1 `meshbay_common` — `protocol.py`

| Change | Detail |
|---|---|
| `IndexEntry` dataclass | Add `hash_version: int = 1` field |
| `index_entry_wire()` | Add `"hash_version": e.hash_version` to the wire dict |

### 2.2 `meshbay_node` — `indexer/indexer.py`

| Change | Detail |
|---|---|
| Constants | `_PARTIAL_THRESHOLD = 40 * 1024 * 1024`, `_PARTIAL_HEAD = 20 * 1024 * 1024`, `_PARTIAL_TAIL = 20 * 1024 * 1024`, `_PARTIAL_MID = 5 * 1024 * 1024` |
| `_scan_file()` | After stat, if `size > _PARTIAL_THRESHOLD`: use partial-read blake3. Set `hash_version=2` on the returned `IndexEntry`. Otherwise: unchanged (full read, `hash_version=1`) |
| `_hash_or_cached()` | Compute expected `hash_version` from file size. Pass it to cache `lookup()`. Store it in cache `put()`. Set it on the returned `IndexEntry` |

**`_scan_file` partial-read implementation:**

```python
def _partial_hash(file_path: Path, size: int) -> str:
    hasher = blake3.blake3()
    with open(long_path(file_path), "rb") as f:
        _feed(hasher, f, _PARTIAL_HEAD)
        f.seek(size - _PARTIAL_TAIL)
        _feed(hasher, f, _PARTIAL_TAIL)
        f.seek(size // 2)
        _feed(hasher, f, _PARTIAL_MID)
    return hasher.hexdigest()

def _feed(hasher, f, nbytes: int) -> None:
    remaining = nbytes
    while remaining > 0:
        chunk = f.read(min(_HASH_CHUNK, remaining))
        if not chunk:
            break
        hasher.update(chunk)
        remaining -= len(chunk)
```

### 2.3 `meshbay_node` — `indexer/cache.py`

| Change | Detail |
|---|---|
| Schema | Add `hash_version INTEGER NOT NULL DEFAULT 1` column to `files` table |
| `_SCHEMA` | New databases get the column via `CREATE TABLE` |
| `_MIGRATE_V2` | `ALTER TABLE files ADD COLUMN hash_version INTEGER NOT NULL DEFAULT 1` — applied on `open()` if the column does not exist |
| `lookup()` | Add `hash_version` parameter. WHERE clause becomes `path = ? AND size = ? AND mtime = ? AND hash_version = ?` |
| `put()` | Add `hash_version` parameter. INSERT includes `hash_version` |
| `CachedEntry` | Add `hash_version: int` field |

**Auto-migration on open:** `IndexCache.open()` runs the ALTER TABLE inside a try/except
(column already exists → no-op). This way a node upgrade just works — no manual step
needed for the cache.

### 2.4 `meshbay_node` — `indexer/group_index.py`

No code change needed. `serialize()` calls `asdict(e)` which includes `hash_version`.
`deserialize()` calls `IndexEntry(**e)` which uses the default `1` for entries written
before v2.

### 2.5 `meshbay_node` — `transport/wire.py`

No code change needed. `index_sync_message()` and `index_delta_message()` call
`index_entry_wire()` which is updated in §2.1.

### 2.6 Client side (browser / desktop)

**No code changes required.** Entries are JavaScript objects; the extra `hash_version`
field is carried through without needing explicit handling:

- `transport.js` — `_applyIndexMessage()` passes the opened payload through. `hash_version`
  rides along on each entry object.
- `group-page.js` — `applyIndex()` / `applyIndexDelta()` store entries as-is in state
  and in IndexedDB.
- `hub-client.js` — IndexedDB cache stores entry objects verbatim.
- `source-merge.js` — merges by `id`. Different hashes naturally don't merge.
- `search-page.js` — aggregates entries from cached indexes. No change.
- `files-app.js` — renders entries. `hash_version` is ignored.

### 2.7 Hub side

**No code changes required.** `swarm_sources.content_hash`, `content_reports.content_hash`,
`content_blocklist.content_hash` are opaque `String(64)` columns. They store whatever
blake3 hex the node provides. No hub migration needed.

### 2.8 Tests

| Test | What it verifies |
|---|---|
| `test_indexer.py` — new cases | Partial hash for file > 40 MB produces `hash_version=2`. File <= 40 MB produces `hash_version=1`. Partial hash is deterministic. Partial hash differs from full hash for the same large file |
| `test_indexer.py` — existing cases | All existing tests still pass (small files, type detection, enrichment, etc.) |
| `test_index_cache.py` — new cases | Cache lookup with `hash_version` match. Cache miss when `hash_version` differs. Schema migration from v1 cache |
| `test_index_cache.py` — existing cases | Unchanged behaviour for v1 entries |
| `test_indexer.py` — round-trip | `GroupIndex.serialize()` → `deserialize()` preserves `hash_version` on entries |
| `test_index_seal_client.py` | Wire format includes `hash_version`, old entries without it deserialize as v1 |

---

## 3. Migration

### 3.1 What actually needs migrating

The node has two relevant stores:

| Store | Location | Content | Migration |
|---|---|---|---|
| `index_cache.db` | `data_dir/index_cache.db` | `(path, size, mtime) → hash` accelerator | Add `hash_version` column |
| In-memory `GroupIndex` | rebuilt from disk on every startup | Current file listing | No migration — rebuilt on next scan |

**The IndexCache is the only persistent store that needs a schema change.** The GroupIndex
is rebuilt by scanning the filesystem on every daemon start. Once the code uses v2 hashing,
the next startup produces v2 hashes for large files automatically.

**The cache auto-migrates.** `IndexCache.open()` adds the `hash_version` column if missing.
Existing rows get `DEFAULT 1`. When the v2 indexer looks up a large file with
`hash_version=2`, the cached v1 entry won't match (different hash_version in WHERE), so
the file is re-hashed with the partial algorithm and the new entry is written with
`hash_version=2`.

This means: **large files are re-hashed lazily on first scan after upgrade.** The first
scan after upgrading to v2 re-reads 45 MB per large file instead of the full content —
already much faster than v1's full read.

### 3.2 Migration script — `QE/migration/migrate_index_v2.sh`

A standalone bash script (not in git — QE/ is gitignored) that:

1. Detects the node's `data_dir` from `node.toml` (default `~/.local/share/meshbay-node/`)
2. Checks that `index_cache.db` exists
3. Runs `ALTER TABLE files ADD COLUMN hash_version INTEGER NOT NULL DEFAULT 1`
4. Optionally (`--purge-large`) deletes cache entries for files > 40 MB, forcing immediate
   re-hash on next scan instead of lazy migration
5. Reports what it did

**Works on Ubuntu and Fedora** — uses only `sqlite3` (present by default on both) and
standard bash.

**Not strictly required** if the code's auto-migration in `IndexCache.open()` is
implemented. The script exists for operators who want to:
- Verify the schema change before restarting
- Force a clean re-hash of all large files in one pass
- Run the migration on a machine where the node isn't installed yet (preparing a data dir)

### 3.3 What the operator does

1. Update the node package (or `pip install -e` in dev)
2. (Optional) Run `QE/migration/migrate_index_v2.sh` to preview or force the migration
3. Restart the node daemon
4. The first scan re-hashes files > 40 MB with partial reads — much faster than before

No client-side action needed. No hub-side action needed.

---

## 4. What does NOT change

- **Files <= 40 MB** — identical hashing, identical `id`, `hash_version=1`.
- **Enrichment** (thumbnails, duration, metadata) — unchanged, still runs after hashing.
- **File downloads** — `file_req` uses the current `id` from the index. After re-indexing,
  clients get the new index with new hashes and request accordingly.
- **Watchdog / reconciliation** — filesystem events trigger the same code paths. New or
  modified files are hashed with the appropriate method based on size.
- **Index encryption / signing** — `GroupIndex.serialize()` and the GEK-sealed wire
  messages are unchanged. `hash_version` rides inside each entry via `asdict()`.
- **Index delta computation** — `GroupIndex.diff()` compares entries by all fields
  (dataclass `__eq__`). A re-indexed file whose hash changed (v1 → v2) appears as a
  deletion of the old id + addition of the new id, which is correct.
- **Hub tables** — opaque content_hash storage, untouched.
- **MNP version** — this is an additive field on index entries. No protocol version bump
  needed. An older node that doesn't send `hash_version` is handled by the default.

---

## 5. Side effects — documented and accepted

1. **Cross-group search**: a file > 40 MB indexed on a v1 node and a v2 node produces
   different `id`s. The search page will not merge them as the same file. The user sees
   two entries instead of one, each from its own group. This resolves itself when both
   nodes upgrade.

2. **First scan after upgrade**: files > 40 MB are re-hashed. With v2 this reads 45 MB
   per file (not the full content), so the re-index is fast — but it is not instant. A
   4 TB library with 1000 large files reads ~44 GB instead of 4 TB.

3. **`content_hash` drift on hub tables**: if a public group's node upgrades, the hashes
   it registers in `swarm_sources` change for large files. Old entries with v1 hashes
   become stale. The swarm registration mechanism's `last_seen` update handles this — stale
   entries age out. No explicit cleanup needed.

4. **Index version bump**: re-indexing sets `GroupIndex.version = int(time.time())`,
   which triggers a full `index_sync` to all connected peers. This is the normal path for
   any index change — it is not new load.

---

## 6. Indexing trigger points — verified safe

| Trigger | Location | Impact of v2 |
|---|---|---|
| Daemon startup — initial scan | `daemon.py:648-674` | Uses `_hash_or_cached()` which applies v2 rules. Safe |
| Create Group wizard — Step 3 | `create-group-page.js:244-248` polls `index-status` | No change — polls progress, doesn't control hashing |
| Settings — add root | `group-settings.js:909-927` | Triggers `retarget()` → scan → `_hash_or_cached()`. Safe |
| Watchdog — file created/modified | `indexer.py:797-816` | Calls `_update_entry()` → `_hash_or_cached()`. Safe |
| Reconciliation loop | `indexer.py:513-650` | Calls `_sweep_available_roots()` → `_hash_or_cached()`. Safe |
| Hot reload — `_reload_config()` | `daemon.py:790-829` | Creates new `DirectoryIndexer` with v2 code. Safe |

---

## 7. Implementation order

1. **`protocol.py`** — add `hash_version` field to `IndexEntry` and `index_entry_wire()`
2. **`cache.py`** — add `hash_version` to schema, auto-migrate on open, update lookup/put
3. **`indexer.py`** — implement `_partial_hash()`, update `_scan_file()` and
   `_hash_or_cached()`
4. **Tests** — new test cases for partial hashing, cache versioning, wire round-trip
5. **`QE/migration/migrate_index_v2.sh`** — standalone migration script
6. **Manual test** — run a node with a mixed library (small + large files), verify:
   - Small files: same hash as before, `hash_version=1`
   - Large files: different hash, `hash_version=2`, 45 MB read
   - Cross-group search: small files merge, large files don't (across v1/v2 nodes)
   - Cache hit on second scan: no re-read
   - Index sync to connected peers: entries carry `hash_version`