<feed xmlns='http://www.w3.org/2005/Atom'>
<title>meshbay.git/packages/meshbay-node/src/meshbay_node/indexer/title_parse.py, branch 0.13</title>
<subtitle>MeshBay — read-only public mirror</subtitle>
<id>https://git.meshbay.org/meshbay.git/atom?h=0.13</id>
<link rel='self' href='https://git.meshbay.org/meshbay.git/atom?h=0.13'/>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/'/>
<updated>2026-08-30T15:05:15Z</updated>
<entry>
<title>chore: scrub copyrighted names from tests, comments and docs</title>
<updated>2026-08-30T15:05:15Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-30T15:05:15Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=d5fc45ff907e618d884a8d2d91c1f2b45e6898d1'/>
<id>urn:sha1:d5fc45ff907e618d884a8d2d91c1f2b45e6898d1</id>
<content type='text'>
Real franchise / show / release-group names had crept back into test
fixtures, code comments, a docstring and docs/mediacenter.md while fixing
the saga-match and misclassification bugs. Replace them all with invented
placeholders ("Some Saga", "A Different Show") and shape descriptions
("a franchise-origin film", "a 3-season show"). Behaviour and assertions
unchanged; 738 node tests still pass.

Record the rule in CLAUDE.md so it stops recurring.

Co-Authored-By: Claude Sonnet 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_018BMLQjqFGCize2KtNBT79v
</content>
</entry>
<entry>
<title>fix(node): stop a numbered saga all matching its first film</title>
<updated>2026-08-30T14:31:34Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-30T14:31:34Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=854b76338072979526ce36fbc176c82f6fc5d0bd'/>
<id>urn:sha1:854b76338072979526ce36fbc176c82f6fc5d0bd</id>
<content type='text'>
Every "&lt;Saga&gt; Episode &lt;N&gt; - &lt;subtitle&gt;" file in a numbered franchise
resolved to the series' first entry (a real, older film). `sequel_variants`
stripped "Episode &lt;N&gt;" and offered the bare "&lt;Saga&gt;" as a candidate query;
that matches the first film's original_title at ratio 1.0 and beat PASS 1's
correct-but-lower hit. A franchise's bare name is very often a real,
different film.

When a Part/Episode/Chapitre/… keyword carried the index, sequel_variants
no longer emits the bare base — only "&lt;base&gt; &lt;digit&gt;" and "&lt;base&gt; &lt;roman&gt;".
Without a keyword ("&lt;Franchise&gt; 3") the bare base is still offered, so that
fix is untouched. Verified live against TMDB: the franchise's episodes each
resolve to their own entry; the earlier numbered-sequel, two-part-film and
franchise-subtitle regressions all hold.

docs/mediacenter.md §10.3.

Co-Authored-By: Claude Sonnet 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_018BMLQjqFGCize2KtNBT79v
</content>
</entry>
<entry>
<title>fix(node): stop a movie with a mangled quality tag being shelved as a series</title>
<updated>2026-08-29T16:45:37Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-29T16:45:37Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=00880fd8b9d362ad391c784f1ee31bd62aa18d62'/>
<id>urn:sha1:00880fd8b9d362ad391c784f1ee31bd62aa18d62</id>
<content type='text'>
Found on demo35: "Some.Film.2017.MULTI.108.grp.mkv" — the release name's
"1080p" truncated to "108" — makes guessit read S01E08, so enrich.py's
flat-library branch (elif ep.episode is not None) filed a standalone film
as a series. "Fix match" then only offered TV results for the phantom
show, so there was no way out from the UI.

- title_parse.has_episode_marker(): true only for an explicit SxxExx /
  1x08 / "Episode N" / "Season N" token, not a bare 3-4 digit run.
- enrich.py: the flat-library branch now needs ep.season AND ep.episode,
  plus either a real marker or the absence of a "(2019)"-style year.
  Every genuine flat-dumped episode in the corpus carries a marker, so
  real shows are untouched; the same misparse on "1280" ("...2013.1280...")
  is covered too.

docs/mediacenter.md §10.2. Known gap left open: no operator control over
the movie/show kind itself — a "this is a movie / a show" toggle would.

Co-Authored-By: Claude Sonnet 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_018BMLQjqFGCize2KtNBT79v
</content>
</entry>
<entry>
<title>feat(node): V8–V11 — show-branch ladder, year-aware _best_match, wider sequel_variants</title>
<updated>2026-08-29T13:43:38Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-29T13:43:38Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=5d28d0c96cc8489645178b483835489779cb3887'/>
<id>urn:sha1:5d28d0c96cc8489645178b483835489779cb3887</id>
<content type='text'>
V8: the TV/show branch of _tmdb_search used the old "first candidate over
0.6 wins" shape. It now shares one _tmdb_ladder helper with the movie
branch — score every candidate query, keep the best, fast-path a
confident primary hit. A year lifted off the show's folder name
(title_parse.year_in, e.g. "Some.Show.2022.S01") rescues a sub-0.6 hit
that lands on the exact year. title_parse.clean_query de-dots a
folder-derived title without naive_title's extension-stripping trap.

V9: _best_match gains an optional `year`. When the top result is not a
confident textual hit (ratio &lt; 0.6) and a year was requested, a
different result of that exact release year is preferred — TMDB already
year-filtered the search, so this is a hard corroboration, not the fuzzy
re-rank §3.3 warns against. A confident top hit is never overridden.
search_movie/search_tv forward the year.

V10: sequel_variants widened — trailing Roman→digit as well as
digit→Roman, spelled-out indices (one..twelve / un..douze / ordinals),
and a "Part N" / "Chapitre N" wrapper. Still empty for a trailing word
that is not an index or a 4-digit year.

V11: when the primary hit is already decent (&gt;= 0.6) and there is nothing
more specific to try (no alternative_title, no sequel variant — only a
punctuation restatement left), the ladder returns without the extra
requests. The clean-title common case is back to one call.

Co-Authored-By: Claude Sonnet 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_018BMLQjqFGCize2KtNBT79v
</content>
</entry>
<entry>
<title>fix(node): correct TMDB movie matching, per-file overrides, rematch</title>
<updated>2026-08-29T12:58:58Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-29T12:58:30Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=71b7a310ce938f072fe20f27eeeadd40685f1ad1'/>
<id>urn:sha1:71b7a310ce938f072fe20f27eeeadd40685f1ad1</id>
<content type='text'>
A batch of wrong poster-grid matches found live on a real library
(2026-08-29): a two-volume film's second part matched the first; a
numbered sequel matched a same-year making-of documentary; several
entries of one franchise matched a single early entry whose localized
TMDB title is the franchise name; one matched nothing. One mechanism:
_tmdb_search returned the first candidate query whose title-similarity
ratio merely cleared 0.6, before alternative_title / the Roman-numeral
variant was ever tried.

Matching:
- title_parse: fold guessit's volume/part number back into display_title
  so the parts of a multi-part film stay distinct in the query, the card
  and the override.
- _tmdb_search: keep a strong PASS 1 fast path (ratio &gt;= 0.85, one
  request), otherwise score every candidate query and pick the best. A
  year-exact rescue lifts a sub-0.6 top hit to the confidence floor only
  when TMDB's own year-filtered result lands exactly on the filename's
  year. No local re-ranking of any single result list; no tmdb.py change.

Fix match / rematch:
- _admin_exec_tmdb_override: a movie override touches its own file only
  (guessit gives a whole franchise one display_title); a show override
  still fans out. Corrected files are marked in media_cache.tmdb_override.
- media_cache: tmdb_override table; clear_file_tmdb / clear_tmdb_matches
  drop auto-resolved matches while sparing manual corrections.
- ops.rematch_video + `meshbay-node video rematch` (loopback endpoint +
  CLI verb): re-resolve a group's video matches after a matcher fix.
  file_tmdb is keyed by content hash and otherwise only pruned on
  deletion, so nothing dislodged a cached match before.
- a rename now drops the stale auto match too (daemon
  _reenrich_renamed_video_entries).

UI:
- VideoDetailModal shows the source filename and resolved TMDB id; an
  unmatched poster gets a badge (3 new video.* i18n keys x 10 locales).
  So a wrong match can actually be identified before hitting Fix match.

docs/mediacenter.md 10.1 records this and the V8-V13 follow-up backlog
(show-branch ladder, year-aware _best_match, wider sequel_variants, the
0.6-0.85 extra calls, movie grid merge, per-card rematch).

Co-Authored-By: Claude Sonnet 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_018BMLQjqFGCize2KtNBT79v
</content>
</entry>
<entry>
<title>fix(video): recognize "S1"/"S2"-style season folders, not just full words</title>
<updated>2026-08-26T14:24:18Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-26T14:24:18Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=bc94dc672d142e989fa4f483f643d4b43f66b02f'/>
<id>urn:sha1:bc94dc672d142e989fa4f483f643d4b43f66b02f</id>
<content type='text'>
Found live: a real show organized its first season as "House.of.the.Dragon.
S01E01...mkv" inside a folder named "S1" (the show name embedded in every
filename, so grouping worked by guessit's own title alone), but its later
seasons as "S02E01. Episode's Own Title.mkv" inside "S2"/"S3" — no show
name in any filename at all, relying entirely on the folder. Those seasons
showed up as loose individual entries instead of grouped under the show:
season_from_folder_name only recognized "season"/"saison"/"livre" as full
words, so "S2" didn't register as a season folder at all, and the
ancestor-based show-grouping fix (previous commits) never triggered for
those seasons.

A bare "S" + 1-2 digits as the *whole* folder name is now recognized too —
anchored to the entire name so it can't match some unrelated folder that
merely starts with "s" followed by digits.
</content>
</entry>
<entry>
<title>fix(video): fix two remaining bugs in Specials numbering and long-season episodes</title>
<updated>2026-08-26T14:04:00Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-26T14:04:00Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=c81eaa0bd07e360ec03406dfc63267c3013a0319'/>
<id>urn:sha1:c81eaa0bd07e360ec03406dfc63267c3013a0319</id>
<content type='text'>
1. A season spanning more than one folder (a per-book Bonus folder nested
   inside every numbered season) got the same synthetic episode numbers
   handed out again in each folder independently — three unrelated
   Specials all showing up as "S0E01". _synthetic_episode_number now ranks
   across the whole show for the target season, not just one file's own
   folder.

2. guessit reads a bare 3-digit leading episode number as a concatenated
   season+episode guess rather than a plain episode number — "100" parses
   as season=1, episode=0, not episode=100, with nothing in its output
   distinguishing that from a real 2-digit episode. A season-like ancestor
   already overrides guessit's season (previous commit); this applies the
   same fix to episode via a direct regex on the leading number, capped at
   3 digits so a leading year (4 digits) is never misread the same way.

Both confirmed against the real library this whole fix was found on:
per-book Bonus features across two books no longer collide, and a real
100th-episode file now resolves to episode 100 instead of 0.
</content>
</entry>
<entry>
<title>fix(video): group every episode under the show's own folder, not a per-file guessit title</title>
<updated>2026-08-26T13:51:10Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-26T13:51:10Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=b1543faad87fd865f8746711ab5f10ac4f331e8b'/>
<id>urn:sha1:b1543faad87fd865f8746711ab5f10ac4f331e8b</id>
<content type='text'>
The mechanism itself was wrong, not just the TMDB matching: a bare episode
numbering convention with no show name in the filename at all
(`001    Episode's Own Title.ext`, no SxxExx, no show prefix — entirely
ordinary on its own) makes guessit invent a "title" from whatever text
follows the number. That text is the individual episode's own name, and
differs for every episode in the folder — trusting it, as the code did,
groups nothing together at all: every episode became its own single-
episode "show", searched against TMDB by that one-off title alone.

Whenever a season-like ancestor folder exists (a numbered season, or
Specials/Bonus/Extras -&gt; season 0), its own root folder now names the show
unconditionally — never a per-file guessit title, which cannot tell a
show's real name from an individual episode's one-off name when the
filename carries no reliable ShowName/SxxExx structure. The walk continues
past *every* consecutive season-like ancestor, not just the first: a
per-season Bonus folder (Show/Season N/Bonus/file.ext) is nested two
levels inside the show, both "Bonus" and "Season N" season-like on their
own, and stopping at the first would hand back "Season N" as the show's
name instead of "Show".

Also adds "livre" ("book") to the season-word vocabulary (§3.4) — some
shows number their seasons that way (Roman numerals) rather than
"season"/"saison".

Supersedes _title_from_show_siblings from the previous commit (removed):
that fallback assumed a per-file title could still be trusted often
enough to be worth borrowing from a sibling season folder — this fix
means it never needs trusting in the first place once a season-like
ancestor exists.

Verified directly against the real library this was found on, not just
synthetic tests: every sample file (a numbered episode, a Book/Bonus
feature, a Book-root "making of") now resolves to the show's real name.
</content>
</entry>
<entry>
<title>fix(node): stop inventing a fake artist from the shared root's name</title>
<updated>2026-08-24T16:30:00Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-24T16:30:00Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=584730f9486e153c6c477293f03a95e78f5ab5f3'/>
<id>urn:sha1:584730f9486e153c6c477293f03a95e78f5ab5f3</id>
<content type='text'>
Music grouping was measured against a real ~5700-file library and came
back worse than a plain file listing. Root cause: the artist/album
ancestor walk always climbed exactly two levels (parent = album,
grandparent = artist) with no idea where the group's own shared root
was. Any file in a flat top-level folder — common here: bare
`Artist/track.mp3`, no album subfolder at all — had its "grandparent"
resolve to the root directory's own name, so the artist got replaced
by the share's name. Measured: 289 of 5664 tracks across 41 real,
unrelated artists (Ben Harper, Dire Straits, Jimi Hendrix, Janis
Joplin, ...) collapsed into one fake artist this way — the single
biggest bucket in the whole library, ahead of every real one.

- `_artist_album_from_ancestors` now takes the entry's own root
  boundary (daemon.py resolves it via `RootSet.split`) and refuses to
  read it as a name. A file sitting in a top-level folder — genuinely
  ambiguous, artist or a standalone album/compilation — is handled by
  `_split_top_level_folder`: split on "Artist - Album" when the
  (cleaned) folder name has that shape, otherwise the whole name
  becomes the artist alone, the more common real case here.
- `_clean_tag` treats known tagger placeholders ("No Artist", a French
  tool's "Nouvel artiste (334)") as absent rather than a real value —
  they were just as truthy as a real name and were locking out the
  fallback that would have done better. "Various Artists" is kept, a
  real compilation credit rather than a placeholder.
- A `title` tag that's the bare filename copied verbatim (track number
  included — found live on a whole CD-single) is stripped through the
  same prefix rule the filename parser already used
  (`title_parse.strip_track_prefix`), since a tag normally wins over
  the parsed title.
- Cover art: only 11% of a 400-file sample had embedded art (expected
  for this era of rip), but 267 loose cover images sit beside the
  tracks across the library (Windows Media Player's `Folder.jpg`/
  `AlbumArt_{guid}_*.jpg`, manual `cover.jpg`) and were never looked
  at. `_find_sibling_cover` checks the track's own folder before
  giving up — measured coverage 11% -&gt; 26% on the same library, zero
  network calls.

`_enriched_attempted` is in-memory and resets on restart, so a node
restart is enough to re-run enrichment over an already-scanned library
with the fixed logic — no rescan flag, no cache to clear by hand.

22 tests in test_enrich_audio.py (11 new), including the exact
regression case end to end through the real pool. Full suite: 1129
passed, no regressions.

Co-Authored-By: Claude Sonnet 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_01KBi7ALLGfwcjBXt57yNMcy
</content>
</entry>
<entry>
<title>feat(node): Music app node-side — indexing, MusicBrainz enrichment, protocol</title>
<updated>2026-08-24T15:12:36Z</updated>
<author>
<name>Christophe Besson</name>
<email>cbesson@gmail.com</email>
</author>
<published>2026-08-24T15:12:36Z</published>
<link rel='alternate' type='text/html' href='https://git.meshbay.org/meshbay.git/commit/?id=941d1a135dd7b03834576855e8e9fdaa24c4e406'/>
<id>urn:sha1:941d1a135dd7b03834576855e8e9fdaa24c4e406</id>
<content type='text'>
Implements the node half of docs/musicbay.md against MNP 0.8:

- IndexEntry gains artist/album/track_no (reuses duration/thumb_hash/
  display_title, already generic). New musicbrainz_config/_enabled and
  music_meta_req/_resp message pairs, mirroring the TMDB shape.
- title_parse.parse_track_filename: track-number-prefix + title parsing,
  fallback-only (embedded tags are the primary source, unlike Videos).
- indexer.enrich_audio.AudioEnricher: mutagen-based tag/embedded-cover
  extraction through its own bounded pool (asyncio.to_thread, no
  subprocess — no ffmpeg-shaped deadlock risk). Gated on "music" in a
  group's enabled_apps rather than a video_root-style scoped folder.
- musicbrainz.py: MusicBrainzClient — no API key (unlike TMDB), just a
  self-imposed ~1 req/s pace and a configurable, non-default User-Agent
  contact string; inert (no calls at all) when no contact is configured,
  never sends an unidentified client.
- media_cache.py: file_mbid/mbid_meta tables alongside the existing TMDB
  ones, cover art reusing the thumbs table via a synthetic
  musicbrainz:{mbid} id, pruned on file deletion.
- roster.py/ops.py/webrtc_server.py: musicbrainz_contact (node-wide) and
  musicbrainz_enabled (per-group, from the start) as signed operator
  settings, ALLOWED_APPS gains "music", _do_music_meta_request resolves
  and caches a release-level MusicBrainz match per (artist, album).
- daemon.py: AudioEnricher/MusicBrainzClient wired alongside the video
  ones; a group's existing library is swept when "music" is newly
  enabled (no video_root equivalent — see musicbay.md §2.1).

41 new tests (musicbrainz.py against a mocked transport, admin-op policy
for both new settings, media_cache round-trip/pruning, enrich_audio
end-to-end against real ffmpeg-generated MP3s). Full suite (common +
node + hub): 1116 passed, no regressions.

Client-side (music-app.js, persistent player bar) not started yet.

Co-Authored-By: Claude Sonnet 5 &lt;noreply@anthropic.com&gt;
Claude-Session: https://claude.ai/code/session_01KBi7ALLGfwcjBXt57yNMcy
</content>
</entry>
</feed>
