<feed xmlns='http://www.w3.org/2005/Atom'>
<title>meshbay.git/packages/meshbay-node/src/meshbay_node/indexer, branch 0.8</title>
<subtitle>MeshBay — read-only public mirror</subtitle>
<id>https://git.meshbay.org/meshbay.git/atom?h=0.8</id>
<link rel='self' href='https://git.meshbay.org/meshbay.git/atom?h=0.8'/>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/'/>
<updated>2026-08-26T16:09:44Z</updated>
<entry>
<title>fix(music): stop sharing one folder's cover across unrelated tracks, and merge various-artists compilations into one album</title>
<updated>2026-08-26T16:09:44Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-26T16:09:44Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=6fb948045b9c563ce591da289cdac1df6bed360c'/>
<id>urn:sha1:6fb948045b9c563ce591da289cdac1df6bed360c</id>
<content type='text'>
Two real, confirmed bugs in a large flat music library:

- enrich_audio.py's sibling-cover fallback assumed one folder is one
  release. A large flat "chart ranking" folder mixing dozens of unrelated
  artists carried several distinct WMP AlbumArt-cache guids (one per
  original album a track was ripped from), and the fallback picked
  whichever one WMP had copied to Folder.jpg — attaching one unrelated
  release's cover to every other track in the folder. Now refuses to pick
  a cover at all once 2+ distinct guids show up, rather than guess.

- music-app.js's groupMusicEntries grouped by artist first, album second,
  so a various-artists compilation (many genuinely different per-track
  artists, one shared album tag, no album-artist tag at all — a real
  ~20-track soundtrack rip has exactly this shape) could never be
  recognized as one release: every track landed alone in its own artist's
  bucket and got folded into a singleton pile. Now detects an album key
  shared across 2+ distinct artist keys and merges those tracks into one
  compilation card under a "Various" heading instead.

Both verified against real, previously-affected files and live in the
browser: the shared wrong cover is gone, and the compilation renders as
one card with all its tracks in order.
</content>
</entry>
<entry>
<title>fix(video): recognize "S1"/"S2"-style season folders, not just full words</title>
<updated>2026-08-26T14:24:18Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-26T14:24:18Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=bc94dc672d142e989fa4f483f643d4b43f66b02f'/>
<id>urn:sha1:bc94dc672d142e989fa4f483f643d4b43f66b02f</id>
<content type='text'>
Found live: a real show organized its first season as "House.of.the.Dragon.
S01E01...mkv" inside a folder named "S1" (the show name embedded in every
filename, so grouping worked by guessit's own title alone), but its later
seasons as "S02E01. Episode's Own Title.mkv" inside "S2"/"S3" — no show
name in any filename at all, relying entirely on the folder. Those seasons
showed up as loose individual entries instead of grouped under the show:
season_from_folder_name only recognized "season"/"saison"/"livre" as full
words, so "S2" didn't register as a season folder at all, and the
ancestor-based show-grouping fix (previous commits) never triggered for
those seasons.

A bare "S" + 1-2 digits as the *whole* folder name is now recognized too —
anchored to the entire name so it can't match some unrelated folder that
merely starts with "s" followed by digits.
</content>
</entry>
<entry>
<title>fix(video): fix two remaining bugs in Specials numbering and long-season episodes</title>
<updated>2026-08-26T14:04:00Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-26T14:04:00Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=c81eaa0bd07e360ec03406dfc63267c3013a0319'/>
<id>urn:sha1:c81eaa0bd07e360ec03406dfc63267c3013a0319</id>
<content type='text'>
1. A season spanning more than one folder (a per-book Bonus folder nested
   inside every numbered season) got the same synthetic episode numbers
   handed out again in each folder independently — three unrelated
   Specials all showing up as "S0E01". _synthetic_episode_number now ranks
   across the whole show for the target season, not just one file's own
   folder.

2. guessit reads a bare 3-digit leading episode number as a concatenated
   season+episode guess rather than a plain episode number — "100" parses
   as season=1, episode=0, not episode=100, with nothing in its output
   distinguishing that from a real 2-digit episode. A season-like ancestor
   already overrides guessit's season (previous commit); this applies the
   same fix to episode via a direct regex on the leading number, capped at
   3 digits so a leading year (4 digits) is never misread the same way.

Both confirmed against the real library this whole fix was found on:
per-book Bonus features across two books no longer collide, and a real
100th-episode file now resolves to episode 100 instead of 0.
</content>
</entry>
<entry>
<title>fix(video): group every episode under the show's own folder, not a per-file guessit title</title>
<updated>2026-08-26T13:51:10Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-26T13:51:10Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=b1543faad87fd865f8746711ab5f10ac4f331e8b'/>
<id>urn:sha1:b1543faad87fd865f8746711ab5f10ac4f331e8b</id>
<content type='text'>
The mechanism itself was wrong, not just the TMDB matching: a bare episode
numbering convention with no show name in the filename at all
(`001    Episode's Own Title.ext`, no SxxExx, no show prefix — entirely
ordinary on its own) makes guessit invent a "title" from whatever text
follows the number. That text is the individual episode's own name, and
differs for every episode in the folder — trusting it, as the code did,
groups nothing together at all: every episode became its own single-
episode "show", searched against TMDB by that one-off title alone.

Whenever a season-like ancestor folder exists (a numbered season, or
Specials/Bonus/Extras -&gt; season 0), its own root folder now names the show
unconditionally — never a per-file guessit title, which cannot tell a
show's real name from an individual episode's one-off name when the
filename carries no reliable ShowName/SxxExx structure. The walk continues
past *every* consecutive season-like ancestor, not just the first: a
per-season Bonus folder (Show/Season N/Bonus/file.ext) is nested two
levels inside the show, both "Bonus" and "Season N" season-like on their
own, and stopping at the first would hand back "Season N" as the show's
name instead of "Show".

Also adds "livre" ("book") to the season-word vocabulary (§3.4) — some
shows number their seasons that way (Roman numerals) rather than
"season"/"saison".

Supersedes _title_from_show_siblings from the previous commit (removed):
that fallback assumed a per-file title could still be trusted often
enough to be worth borrowing from a sibling season folder — this fix
means it never needs trusting in the first place once a season-like
ancestor exists.

Verified directly against the real library this was found on, not just
synthetic tests: every sample file (a numbered episode, a Book/Bonus
feature, a Book-root "making of") now resolves to the show's real name.
</content>
</entry>
<entry>
<title>fix(video): a Specials/Bonus folder's episodes were matched to TMDB as unrelated standalone movies</title>
<updated>2026-08-26T13:29:05Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-26T13:22:17Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=4c45792c861c47d44ebdbbab4130e37fba5d6795'/>
<id>urn:sha1:4c45792c861c47d44ebdbbab4130e37fba5d6795</id>
<content type='text'>
Found live: a "Specials" folder full of one-off-named bonus episodes had
every file appear as its own poster, matched against TMDB by its own
title, because guessit finds no season/episode grammar at all in a
filename with no SxxExx of its own — so the classification (episode vs
standalone movie), based solely on that, fell to the movie branch. Real,
unrelated films happened to share several of those one-off titles and
matched confidently, one per Special, cluttering the Videos view with
dozens of wrong posters instead of grouping under the show.

An ancestor folder saying this is part of a show — a numbered season, or
Specials/Bonus/Extras -&gt; season 0 — is now trusted over the filename
having no SxxExx of its own. The show's name can't come from this file's
own guessit title (that's the bug) or from siblings in the same Specials
folder (every one of them has the same gap) — it's borrowed from the
show's ordinary season folders next door, which do carry it in the usual
ShowName.SxxExx shape (_title_from_show_siblings). A synthetic, stable
episode number (alphabetical rank among the folder's video files) stands
in for the real one nothing in a Specials folder provides.

Generic, not specific to a folder literally named "Specials": the same
fallback fires for any season-like ancestor folder whose files lack
per-file episode grammar, numbered seasons included (covered by a
dedicated test).

Also hardens _title_from_siblings to require the sibling's own episode
number too, not just a title — otherwise it would borrow one Special's
one-off title as if it were representative, on a folder like this one.

Existing already-indexed entries will need a rescan (restart the node) to
be re-enriched under this logic — nothing re-derives them on its own.
</content>
</entry>
<entry>
<title>fix(node,hub): key music/media metadata lookups by file_id, not path</title>
<updated>2026-08-25T23:08:58Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-25T23:08:58Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=2d144d76cee55cf8faaacf196e716a0930dfd7e9'/>
<id>urn:sha1:2d144d76cee55cf8faaacf196e716a0930dfd7e9</id>
<content type='text'>
IndexEntry.path is the *folder* a file is in (indexer.py's
_virtual_dir docstring: "the directory a file appears in"), not the
file itself. GroupIndex.get_entry_by_path() treated it as if it named
one file, and every one of its four callers did too:
_do_music_meta_request, _do_media_meta_request, _do_tmdb_override, and
_admin_exec_tmdb_override. Any two files sharing a folder — an album is
one folder with many tracks, a season is one folder with many episodes
— collided: a lookup by path silently returned whichever entry the
index happened to iterate to first, regardless of which file the
client actually asked about.

Found live (2026-08-25): three unrelated albums ("High Tone - Various",
two "Le Peuple de l'Herbe" albums) all showed the same MusicBrainz
cover, because all their representative tracks happened to sit in one
"high_tone" folder alongside a track that legitimately matched that
cover. A force-reload didn't help — the bug is server-side, not a
stale client state.

Fixed by keying these four request/response pairs by `file_id` (the
entry's own content hash — already unique, already how every other
lookup in the system identifies a file) instead of `path`, both in the
wire messages (music_meta_req/resp, media_meta_req/resp, tmdb_override)
and in music-app.js/video-app.js's own hooks. GroupIndex.get_entry_by_path
is now unused and removed — GroupIndex.get_entry(file_id) already did
the right thing.

No test previously exercised either handler with two entries sharing a
folder — the only existing coverage (test_tmdb_override_policy.py) gave
each entry its own folder, so the bug never had a chance to show up.
Added that scenario there and in two new test files, all confirmed
failing against the pre-fix code before being confirmed green against
the fix.

Co-Authored-By: Claude Sonnet 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_013XSohfUQQiaE77qyFLgSv3
</content>
</entry>
<entry>
<title>feat(node): share the (path,size,mtime)-&gt;hash index cache across every group</title>
<updated>2026-08-25T22:40:20Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-25T22:40:20Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=37d8d9c15c982f2da17b2fad4ea1a90613b560a6'/>
<id>urn:sha1:37d8d9c15c982f2da17b2fad4ea1a90613b560a6</id>
<content type='text'>
An operator routinely shares the same physical folder into more than one
group (a music library, a Séries drive) — IndexCache used to be opened
once per group (data_dir/{group_id}/index_cache.db), so the second group
to reference an already-fully-hashed multi-terabyte folder paid the same
full content read the first one did. IndexCache itself carried no
group_id in its schema; only daemon.py's wiring did. Now one instance,
opened once at startup (data_dir/index_cache.db), shared by every group's
DirectoryIndexer.

Confirmed against a real deployment (2026-08-25/26): a group sharing an
already-indexed folder with an existing group indexes it instantly, with
zero rehashing.

Also fixes a related cross-group correctness gap found during this work:
media_cache.db (thumbnails, TMDB/MusicBrainz metadata — already node-wide,
untouched by this change) was pruned for a file the moment it left *one*
group's index, even if another group's index still held the same content
hash — forcing a redundant re-fetch/re-probe/re-thumbnail for a group that
never actually lost anything. Prune now runs only once no group's index
references the file_id any more.

Adds a node admin UI action ("Maintenance" card, prune-index-cache) to
drop cache rows that no longer belong to any group's roots — skips
anything under a root that is merely temporarily unavailable (indexer.py's
"a root that goes away freezes, never empties" rule extends to this
cache too, or a reconnected drive would pay a full rehash for no reason).

Co-Authored-By: Claude Sonnet 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_013XSohfUQQiaE77qyFLgSv3
</content>
</entry>
<entry>
<title>fix(hub,node): Create Group wizard silently skipped apps, and lost track of scanning progress</title>
<updated>2026-08-25T16:18:48Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-25T16:18:48Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=3c44f55f6b0aba77c7ad57d0a5ebe3e55b409473'/>
<id>urn:sha1:3c44f55f6b0aba77c7ad57d0a5ebe3e55b409473</id>
<content type='text'>
Two real-world bugs found together while testing multi-root group
creation:

- CreateGroupWizard only sent the enabled-apps PUT when the operator had
  *unchecked* something, assuming "every box left checked" already matched
  the node's own default (Roster.DEFAULT_APPS = chat, files). It doesn't —
  so leaving every app checked, the common case, silently left
  Videos/Music/Photos disabled on the node. Now sent unconditionally.

- The wizard's "add extra roots" step never polled index-status, so once
  step 3 (which only watches the first/upload root) finished, the
  progress bar froze while the node kept scanning the remaining roots for
  minutes, unwatched. Added waitForRootsIndexed (platform.js), mirroring
  waitForGroupHosted's own race handling.

That fix exposed a deeper one: indexer.py's _scan_root() only flipped
`progress.scanning` on *after* walking the directory and stat()-ing every
file — both off-loop, but slow enough on a large root that a poller's
grace period (waitForRootsIndexed's 5s) could expire before ever
observing `scanning: true` (confirmed against production logs: a GEK-init
step fired 5.058s after a root started scanning, matching the grace
period almost exactly). The stat() pass was also a synchronous loop
directly on the asyncio event loop — blocking the whole daemon (WebRTC,
chat, admin UI) for as long as it took on a root with many files. Both
fixed: `scanning` now flips on before the walk starts, and stat()-ing is
now off-loop too (_size_files).

Co-Authored-By: Claude Sonnet 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_013XSohfUQQiaE77qyFLgSv3
</content>
</entry>
<entry>
<title>fix(node): make Photos/Video/Music enrichment survive node restarts</title>
<updated>2026-08-25T10:52:32Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-25T10:52:32Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=9e66c11b15b7103e963dcecf88fb152ed1e74253'/>
<id>urn:sha1:9e66c11b15b7103e963dcecf88fb152ed1e74253</id>
<content type='text'>
Only thumbnail bytes were ever durable in media_cache.db — every other
derived field (photo width/height/EXIF, video ffprobe duration/dims,
audio cover art) lived solely on the in-memory GroupIndex entry, so a
node restart re-decoded every photo through Pillow, re-ran ffprobe on
every video, and re-scanned for every album cover from scratch, even
though the answers already sat in the cache.

Adds photo_meta and video_meta tables (content-only fields, keyed by
file_id) and checks them before doing the expensive work. Audio gets no
new table: mutagen reads tags and duration in one inseparable call, so
caching duration alone buys nothing — instead cover-art extraction alone
is skipped via a new skip_cover flag when a cached cover already exists.

Deliberately excluded from all three caches: anything derived from the
filename or folder path (video display_title/season/episode via guessit,
audio artist/album folder-fallback) — those must keep being recomputed
fresh so a rename/move is still correctly re-derived by the existing
_reenrich_renamed_*_entries mechanisms, instead of silently handing back
a stale parse under the new name/location.

Regression tests prove cache reuse by deleting the source file (or cover)
between two enrichment runs, and prove rename/move correctness survives
the new cache by renaming/moving to a path that never exists on disk.
</content>
</entry>
<entry>
<title>feat: add Photos group app</title>
<updated>2026-08-25T09:46:17Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-25T09:46:17Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=2fcdd07d1e5d331ad02b723f1c45603a0989c264'/>
<id>urn:sha1:2fcdd07d1e5d331ad02b723f1c45603a0989c264</id>
<content type='text'>
A new group application (docs/apps.md's plug-in mechanism), following the
plan in docs/photos.md. Unlike Videos/Music: several photo roots per group
instead of one (photo_roots is a set, one signed op replaces it whole),
a single album-grid view with no third-party matching step, and per-photo
info read from the file's own EXIF at index time — no metadata service,
no credential, no outbound network call at all.

Protocol (meshbay-common, MNP 0.10 -&gt; 0.11, additive): `taken_at`/`camera`
on IndexEntry; `photo_roots`/`photo_roots_ack`; `OP_PHOTO_ROOTS`.

Node: roster.py stores photo_roots as a group_settings entry (JSON list,
same shape as enabled_apps); ops.py/webrtc_server.py validate and sign the
whole set in one op, same pattern as apps_enabled; a new PhotoEnricher
(indexer/enrich_photo.py) runs Pillow in its own small bounded pool,
separate from the video/audio pools, producing a resized thumbnail plus
the two EXIF fields — never GPS, checked by a grep-based regression test.

Client: photos-app.js — one album card per directory containing images,
a per-album photo grid, and a lightbox with next/previous (keyboard and
buttons), zoom in/out/fit/100% starting from the actual on-screen fit
percentage, and a "zip this album" button reusing files-app.js's own zip
mechanism (lifted into file-utils.js's downloadDirectory so both call the
same implementation). group-settings.js gets an add/remove multi-root
picker, distinct from Videos/Music's single-value one.

Bugs found and fixed before this ever shipped, worth keeping the story of:

- enrich_photo.py read width/height from the raw image *before* applying
  EXIF orientation correction, and read DateTimeOriginal off the plain
  0th-IFD Exif object — a real camera stores it in the Exif sub-IFD, which
  Pillow only exposes via get_ifd(Exif). A flat, hand-built EXIF dict
  round-trips through Pillow either way, which is exactly what would have
  hidden both bugs; the regression test builds EXIF with piexif instead,
  matching what real hardware produces.
- photos-app.js's album grouping stripped a trailing path segment from
  entry.path under the assumption it still carried a filename — it
  doesn't (files-app.js's own convention: e.path is already the
  containing directory), so every album collapsed one level into its
  parent. Found live against a real multi-folder library.
- transport.js's ADMIN_OP_TYPES allowlist (already the fix for an
  identical bug on video_root/apps_enabled, see 4783d81) was missing
  photo_roots: its admin_challenge matched no pending request and was
  silently dropped, so saving a photo root just timed out after 30s with
  no error.
- daemon.py pruned a thumbnail when its file left the index (root removed
  or reconfigured) but never forgot the content hash was "already
  attempted" — the same bytes reappearing under a renamed/relocated root
  (an operator's real workflow) were then permanently skipped, forever,
  with nothing to indicate why. Discarding the attempt alongside the
  cache entry on prune is what makes pruning actually reversible.
- packages/meshbay-client's app:// protocol handler served every file
  with no Cache-Control header, so Chromium was free to serve a stale
  cached copy indefinitely — none of several `npm run sync-ui` + reload
  cycles during development actually picked up the new code until the
  renderer's disk cache was cleared by hand. Now sends Cache-Control:
  no-store.
- the lightbox's zoomed image used flex centering (align-items/
  justify-content: center) combined with overflow: auto — a well-known
  trap where the browser centers overflowing content by shifting it, and
  the leading half of that overflow (here, the top of a zoomed photo)
  sits outside what the scrollport can actually reach. Reported live as
  "unusable". Fixed by switching to top/left alignment once zoomed.

Co-Authored-By: Claude Sonnet 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_01TiZG4AuSnxHohQMpwTHTyL
</content>
</entry>
</feed>
