summaryrefslogtreecommitdiffstats
path: root/packages/meshbay-node/src/meshbay_node/indexer/enrich_audio.py
blob: dd6b2c09c2b5e9dc3faeb564988df3e6f37303cb (plain) (blame)
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
"""
Index-time enrichment for the Music group app: embedded tag/cover
extraction (mutagen) and filename-parse fallback for a newly-added audio
IndexEntry (docs/MESHBAY_DESIGN.md §9.8).

Runs through its own small bounded worker pool, the same discipline as the
Videos app's `enrich.py` — separate from any other pool, never blocking a
scan or the watchdog. Unlike `enrich.py`, this one shells out to nothing:
`mutagen` is pure Python, synchronous I/O only, so there is no subprocess to
spawn, no pipe to drain, and no ffmpeg-shaped deadlock risk here at all —
the bounded pool exists to keep a large library's indexing burst bounded,
not to contain a process. Reads run via `asyncio.to_thread` so they never
block the event loop.

MusicBrainz lookups are **not** done here. Tag/cover extraction is free and
local, so it runs for every audio file the Music app is enabled for,
regardless of whether MusicBrainz itself is turned on for the group — the
flat view (docs/MESHBAY_DESIGN.md §9.8) needs nothing more than this. MusicBrainz
is a separate, lazy, per-request enrichment (`music_meta_req`, handled in
webrtc_server.py), the same "fetched on demand, cached once" shape TMDB
already uses.

**Revised 2026-08-24** against a real ~5700-file library (folder-per-artist
mostly, but not uniformly). Two findings drove this revision, both confirmed with
real data before writing the fix:

1. The original ancestor walk always went up two levels (parent = album,
   grandparent = artist) with no idea where the group's shared root itself
   was. For any file sitting in a *top-level* folder — common here: bare
   `Artist/track.mp3`, no album layer at all — the "grandparent" it found
   was the root directory's own name, so every such artist got relabelled
   as an album *of* a fake band named after the share. Measured: 289 of
   5664 tracks (41 real, unrelated artists) collapsed into one bucket this
   way — the single biggest "artist" in the whole library, by a wide
   margin, which is exactly the kind of thing a person notices first and
   loses trust in. Fixed by passing the root's own path down so the walk
   refuses to use it as a name (`_artist_album_from_ancestors` below).
2. A tag being *present* is not the same as it being *meaningful*: some
   taggers write a literal placeholder ("No Artist", a French tool's
   "Nouvel artiste (334)") instead of leaving the field empty, and that
   placeholder is just as truthy as a real name — it was winning over the
   filename/folder fallback that would have done better. `_clean_tag`
   below turns a known placeholder back into "absent" before anything else
   sees it.

WMA and Musepack read their tags outside mutagen's generic "easy" interface
(neither format has an Easy* wrapper — `MutagenFile(path, easy=True)` just
hands back the raw tag object, whose keys are format-specific and don't
match the generic title/artist/album/tracknumber names the easy interface
normally exposes elsewhere). Confirmed against real files before writing
the fix: WMA's real keys are `Title`/`Author`/`WM/AlbumTitle`/
`WM/TrackNumber`, not "artist"/"album"; Musepack's are plain APEv2 keys
(`Title`/`Artist`/`Album`/`Track`). Musepack also can't rely on mutagen's
format auto-detection (`MutagenFile(path)` with no explicit class) — it
misidentified a real .mpc file as MP3 often enough in a spot-check that the
extension is used to pick the class directly instead. Neither format has a
convenient embedded-cover path here, so both fall through to the sibling
image-file fallback for cover art, same as everything else.
"""

import asyncio
import logging
import re
from collections.abc import Awaitable, Callable
from pathlib import Path

import blake3
from meshbay_common.protocol import IndexEntry
from mutagen import File as MutagenFile
from mutagen.asf import ASF
from mutagen.musepack import Musepack

from meshbay_node.indexer import title_parse
from meshbay_node.media_cache import MediaCache

log = logging.getLogger(__name__)

# Higher than the video pool's default (2): mutagen reads a few KB of tag
# data synchronously, no subprocess, no decode — cheap enough that a wider
# pool doesn't cost much and finishes a large library's initial scan sooner.
DEFAULT_MAX_CONCURRENT = 4
READ_TIMEOUT_SECS = 10

_TRACK_NO_RE = re.compile(r"\d+")

# Values some taggers write instead of leaving a field empty — treated as
# "absent" so the fallback chain gets a chance instead of locking in a
# placeholder. Not exhaustive by construction; each entry here was seen in
# a real file, not guessed. "Various Artists" is deliberately *not* here —
# that one is a real, meaningful compilation credit worth keeping as its
# own bucket, not a placeholder. "Album inconnu (<date>)"/standalone
# "Inconnu" is a French ripping tool's own auto-generated placeholder, seen
# on real WMA files — same idea as "Unknown Artist", different language.
_PLACEHOLDER_RE = re.compile(
    r"^(no\s*artist|unknown(\s+artist)?|inconnu(e)?|"
    r"album\s+inconnu\s*\([^)]*\)|"
    r"nouvel(le)?\s+artiste\s*\(\d+\)|"
    r"nouveau\s+titre\s*\(\d+\)|track\s*\d*|<unknown>)$",
    re.IGNORECASE,
)


def _clean_tag(value: str | None) -> str | None:
    if not value:
        return None
    v = value.strip()
    return v if v and not _PLACEHOLDER_RE.match(v) else None


def _extract_cover(mf) -> bytes | None:
    """
    Best-effort embedded cover art across the tag formats mutagen exposes
    differently: ID3 (MP3) keeps pictures as APIC frames on `.tags`, FLAC
    exposes `.pictures` on the file object itself, MP4/M4A keeps a `covr`
    atom on `.tags`. Returns the first picture found, or None — most of a
    real library has no embedded art at all, which is not an error.
    """
    tags = mf.tags
    if tags is not None and hasattr(tags, "getall"):
        pics = tags.getall("APIC")
        if pics:
            return bytes(pics[0].data)
    pictures = getattr(mf, "pictures", None)
    if pictures:
        return bytes(pictures[0].data)
    if tags is not None and hasattr(tags, "get"):
        covr = tags.get("covr")
        if covr:
            return bytes(covr[0])
    return None


# Loose image files sitting beside the tracks — very common for exactly the
# era of rip this library mostly is (Windows Media Player's own
# `Folder.jpg`/`AlbumArt_{guid}_(Large|Small).jpg`, or a manually dropped
# `cover`/`front.jpg`). Measured: 267 such images across one real library,
# none of them ever considered before this revision — embedded-art-only
# missed them entirely. Ranked so a deliberately-named cover file wins over
# WMP's cache thumbnail when both exist in the same folder.
_COVER_STEM_RANK = ("cover", "folder", "front", "albumart")
_COVER_EXTS = (".jpg", ".jpeg", ".png")

# Windows Media Player's per-folder thumbnail cache: one `AlbumArt_{guid}_*`
# pair per *release* it ever cached art for in that folder, keyed by a GUID
# tied to that release — never to a track. Real folder, real GUIDs (module
# docstring's revision note): a ~500-track flat "chart ranking" rip mixing
# dozens of unrelated artists carried seven distinct guids here, each an art
# leftover from one different original album. `Folder.jpg`/`AlbumArtSmall.jpg`
# are WMP's own copy of just *one* arbitrary one of those for the folder icon
# — fine when a folder really is one release (one guid, or none), meaningless
# once two or more show up: there is no way to tell which track it belongs to,
# so no candidate in the folder can be trusted as "the" cover for any of them.
_WMP_ALBUMART_GUID_RE = re.compile(r"^albumart_\{([0-9a-f-]{36})\}_(large|small)$", re.IGNORECASE)


def _find_sibling_cover(folder: Path) -> Path | None:
    try:
        candidates = [p for p in folder.iterdir()
                     if p.is_file() and p.suffix.lower() in _COVER_EXTS]
    except OSError:
        return None
    if not candidates:
        return None

    guids = {m.group(1).lower() for p in candidates
             if (m := _WMP_ALBUMART_GUID_RE.match(p.stem))}
    if len(guids) >= 2:
        return None

    def rank(p: Path) -> int:
        stem = p.stem.lower()
        for i, name in enumerate(_COVER_STEM_RANK):
            if name in stem:
                return i
        return len(_COVER_STEM_RANK)
    candidates.sort(key=lambda p: (rank(p), p.name))
    return candidates[0]


def _first_tag_value(values) -> str | None:
    return str(values[0]) if values else None


def _asf_tags_from_object(asf_file: ASF) -> tuple[dict, float | None]:
    """
    Pure mapping from an already-opened ASF object into our normalized
    dict — split out from the file-open call so the mapping itself can be
    exercised directly. Real ASF key names, confirmed against a real
    library sample and a freshly-encoded fixture alike: `Title`, `Author`
    (not "Artist"), `WM/AlbumTitle`, `WM/TrackNumber`.
    """
    tags: dict = {}
    duration = getattr(asf_file.info, "length", None) if asf_file.info is not None else None
    if asf_file.tags is not None:
        title = _clean_tag(_first_tag_value(asf_file.tags.get("Title")))
        if title:
            tags["title"] = title
        artist = _clean_tag(_first_tag_value(asf_file.tags.get("Author")))
        if artist:
            tags["artist"] = artist
        album = _clean_tag(_first_tag_value(asf_file.tags.get("WM/AlbumTitle")))
        if album:
            tags["album"] = album
        track_raw = _first_tag_value(asf_file.tags.get("WM/TrackNumber"))
        if track_raw:
            m = _TRACK_NO_RE.match(track_raw)
            if m:
                tags["track_no"] = int(m.group())
    return tags, duration


def _musepack_tags_from_object(mpc_file: Musepack) -> tuple[dict, float | None]:
    """
    Pure mapping from an already-opened Musepack object's APEv2 tags into
    our normalized dict — split out the same way as `_asf_tags_from_object`
    (and for the same testability reason: there is no encoder available
    here to produce a real .mpc fixture from tags alone, mutagen included).
    """
    tags: dict = {}
    duration = getattr(mpc_file.info, "length", None) if mpc_file.info is not None else None
    if mpc_file.tags is not None:
        for field, key in (("title", "Title"), ("artist", "Artist"), ("album", "Album")):
            cleaned = _clean_tag(_first_tag_value(mpc_file.tags.get(key)))
            if cleaned:
                tags[field] = cleaned
        track_raw = _first_tag_value(mpc_file.tags.get("Track"))
        if track_raw:
            m = _TRACK_NO_RE.match(track_raw)
            if m:
                tags["track_no"] = int(m.group())
    return tags, duration


def _read_format_specific_tags(path: Path) -> tuple[dict, float | None] | None:
    """
    Returns None when the generic "easy" interface below already covers
    this extension — only WMA and Musepack need a bypass (docstring at the
    top of this module).
    """
    suffix = path.suffix.lower()
    if suffix == ".wma":
        try:
            return _asf_tags_from_object(ASF(str(path)))
        except Exception:
            return {}, None
    if suffix == ".mpc":
        try:
            return _musepack_tags_from_object(Musepack(str(path)))
        except Exception:
            return {}, None
    return None


def _read_generic_tags_and_cover(
    path: Path, skip_cover: bool = False,
) -> tuple[dict, float | None, bytes | None]:
    tags: dict = {}
    duration: float | None = None
    try:
        easy = MutagenFile(str(path), easy=True)
    except Exception:
        easy = None
    if easy is not None:
        if easy.info is not None:
            duration = getattr(easy.info, "length", None)
        for field in ("title", "artist", "album"):
            values = easy.get(field)
            cleaned = _clean_tag(str(values[0])) if values else None
            if cleaned:
                tags[field] = cleaned
        track_raw = easy.get("tracknumber")
        if track_raw:
            m = _TRACK_NO_RE.match(str(track_raw[0]))
            if m:
                tags["track_no"] = int(m.group())

    cover: bytes | None = None
    if not skip_cover:
        try:
            raw = MutagenFile(str(path))
        except Exception:
            raw = None
        if raw is not None:
            try:
                cover = _extract_cover(raw)
            except Exception:
                cover = None

    return tags, duration, cover


def _read_tags_and_cover(
    path: Path, skip_cover: bool = False,
) -> tuple[dict, float | None, bytes | None]:
    """
    Synchronous — always called via asyncio.to_thread. Returns a partial
    `tags` dict (only keys actually found and not a known placeholder:
    title/artist/album/track_no), duration in seconds (None if unreadable),
    and raw cover bytes (None if absent, embedded and sibling-file both
    checked). Never raises for an unreadable/corrupt file — the caller
    falls back to filename parsing entirely in that case.

    `skip_cover` is set when the caller already has a cached cover for this
    exact content (AudioEnricher._run, keyed by the file's own content
    hash) — it skips the second file open plus the sibling-directory scan
    entirely, neither of which the rename-refresh path needs redone: a
    cover is derived from content, never from the filename.
    """
    format_specific = _read_format_specific_tags(path)
    if format_specific is not None:
        tags, duration = format_specific
        cover = None  # neither WMA nor Musepack has a convenient embedded-cover path here
    else:
        tags, duration, cover = _read_generic_tags_and_cover(path, skip_cover=skip_cover)

    if "title" in tags:
        # Some taggers copy the bare filename into `title` verbatim, track
        # number included — a tag normally wins over the filename parse, so
        # that pollution would otherwise beat a cleaner one (docstring above).
        tags["title"] = title_parse.strip_track_prefix(tags["title"]) or tags["title"]

    if cover is None and not skip_cover:
        cover = _read_sibling_cover(path)

    return tags, duration, cover


def _read_sibling_cover(path: Path) -> bytes | None:
    """
    The cover image sitting beside the track, as bytes.

    Its own function because it is the one part of the read that depends on the
    *folder* rather than on the file's bytes: `audio_meta` remembers what the
    bytes said and lets `AudioEnricher._run` skip opening the file at all, and
    this still has to run on top of that, or a cover dropped in after the first
    pass could never be found again. Blocking; called via asyncio.to_thread.
    """
    sibling = _find_sibling_cover(path.parent)
    if sibling is None:
        return None
    try:
        return sibling.read_bytes()
    except OSError:
        return None


# A folder name used as a last-resort artist/album, cleaned of the
# punctuation-as-separator and release-tag noise this era of rip is full of
# (underscores standing in for spaces, a bitrate/quality tag still attached
# to the name). Same spirit as title_parse.naive_title for video, kept
# separate because the junk vocabulary differs (bitrates and rip tags, not
# edition/language tags).
# A whole bracketed/parenthesized group is dropped if it contains any rip-tag
# word ("[MP3 320kbps Album]", "(VBR HQ mp3s)") — matching one keyword at a
# time left the rest of a multi-word group behind ("Album]" survived a first
# version of this that only recognized "full album" as one fixed phrase).
# Bare tokens outside brackets are stripped on their own.
_RIP_TAG_GROUP_RE = re.compile(
    r"[\[\(][^\[\]\(\)]*\b(vbr|cbr|flac|mp3s?|kbps|hq|album)\b[^\[\]\(\)]*[\]\)]",
    re.IGNORECASE,
)
_RIP_TAG_BARE_RE = re.compile(r"\b(vbr|cbr|flac|hq)\b", re.IGNORECASE)


def _clean_folder_name(name: str) -> str:
    cleaned = re.sub(r"[._]+", " ", name)
    cleaned = _RIP_TAG_GROUP_RE.sub(" ", cleaned)
    cleaned = _RIP_TAG_BARE_RE.sub(" ", cleaned)
    cleaned = re.sub(r"\s+", " ", cleaned).strip(" -")
    return cleaned or name


def _split_top_level_folder(name: str) -> tuple[str, str | None]:
    """
    The one ancestor level left once a file's folder turns out to sit
    directly under the group's root (§ below) — there is no further
    ancestor to call "artist" without leaving the root entirely. The common
    convention for a single-release folder at that level is
    "Artist - Album ...junk..." (e.g. a soundtrack folder named after the
    film, or a release folder with the ripper's tags still attached); split
    on the first " - " when the cleaned name has one. Otherwise the whole
    (cleaned) name becomes the artist alone, which is the *more* common real
    shape here — a flat per-artist folder with no album subfolder at all.
    """
    cleaned = _clean_folder_name(name)
    m = re.match(r"^(.{2,60}?)\s*-\s*(.{2,80})$", cleaned)
    if m:
        return m.group(1).strip(), m.group(2).strip()
    return cleaned, None


def _artist_album_from_ancestors(
    file_path: Path, root_path: Path | None,
) -> tuple[str | None, str | None]:
    """
    `Artist/Album/track.mp3` is one real shape in this library, but not the
    only one — measured directly against it (module docstring). This walk
    refuses to climb past `root_path` (the group's own shared directory):
    doing so used to read the root's own name as "the artist", which is
    never true and was the single largest source of bad groupings found.

    Three cases, by how many folders separate the file from the root:
      0 (file sits directly in the root) — no folder context at all.
      1 (a top-level folder) — ambiguous by construction; see
        `_split_top_level_folder`.
      2+ — the classic Artist/Album layout: immediate parent is the album,
        its parent is the artist.

    `root_path` is None when the caller couldn't resolve which named root
    this entry belongs to (should not happen in practice — `entry.path`
    always names one — but the walk still terminates safely at the
    filesystem root rather than looping, same as before this revision).
    """
    leaf = file_path.parent
    if root_path is not None and leaf == root_path:
        return None, None
    parent = leaf.parent
    if root_path is not None and parent == root_path:
        return _split_top_level_folder(leaf.name)
    if leaf == parent:  # filesystem root reached without ever matching root_path
        return None, None
    return _clean_folder_name(parent.name), _clean_folder_name(leaf.name)


class AudioEnricher:
    """Owns the node's bounded audio index-time enrichment pool."""

    def __init__(self, media_cache: MediaCache, max_concurrent: int = DEFAULT_MAX_CONCURRENT):
        self._media_cache = media_cache
        self._sem = asyncio.Semaphore(max_concurrent)
        self._tasks: set[asyncio.Task] = set()

    def spawn(
        self, entry: IndexEntry, file_path: Path,
        on_done: Callable[[str, dict], Awaitable[None]],
        root_path: Path | None = None,
    ) -> asyncio.Task:
        """
        Fire-and-forget one file's enrichment — same contract as
        `enrich.Enricher.spawn`: `on_done(file_id, fields)` is awaited with
        the index fields to merge once ready, never blocks the caller, and
        the returned task must be held by the caller for the same reason
        `WebRTCPeerSession._spawn` holds streaming tasks (a bare
        `ensure_future` can be garbage-collected mid-flight). `root_path` is
        the caller's own root boundary for this entry (daemon.py resolves
        it via `RootSet.split`) — see `_artist_album_from_ancestors`.
        """
        task = asyncio.ensure_future(self._run(entry, file_path, on_done, root_path))
        self._tasks.add(task)

        def _cleanup(t: asyncio.Task) -> None:
            self._tasks.discard(t)
            if not t.cancelled() and t.exception():
                log.error("Audio enrichment failed for %s: %s", entry.id[:12], t.exception(),
                          exc_info=t.exception())
        task.add_done_callback(_cleanup)
        return task

    async def _run(
        self, entry: IndexEntry, file_path: Path,
        on_done: Callable[[str, dict], Awaitable[None]],
        root_path: Path | None,
    ) -> None:
        async with self._sem:
            fields: dict = {}

            # entry.id is the file's own content hash, so everything a read of
            # those bytes produced is cacheable under it — the tags, the
            # duration, and whether the embedded-art scan has already run
            # (media_cache.audio_meta, plus get_thumb_hash_by_file_id for the
            # cover itself). With both in hand the file is not opened at all,
            # which is the whole point: this pass used to re-read every audio
            # file on the node at every start, and until it landed the index
            # the node served carried no artist on any track.
            #
            # What is *not* cached, and must not be: the fallbacks below.
            # artist/album/title/track_no can come from the filename or the
            # folder name (_artist_album_from_ancestors above) when no tag is
            # present, and those a rename has to re-derive —
            # _reenrich_renamed_audio_entries exists for precisely that. The
            # raw tag is not rename-sensitive, so the tag is what is stored and
            # the chain still runs live on top of it. The sibling-image scan is
            # the other one: it reads the folder, not the file, so a cover
            # dropped in afterwards must still be found.
            cached_thumb_hash = await self._media_cache.get_thumb_hash_by_file_id(entry.id)
            cached = await self._media_cache.get_audio_meta(entry.id)
            cover_settled = bool(cached_thumb_hash) or bool(cached and cached["cover_seen"])
            if cached and cover_settled:
                tags = {k: cached[k] for k in ("title", "artist", "album", "track_no")
                        if cached[k] is not None}
                duration = cached["duration"]
                cover = None if cached_thumb_hash else await asyncio.to_thread(
                    _read_sibling_cover, file_path)
            else:
                try:
                    tags, duration, cover = await asyncio.wait_for(
                        asyncio.to_thread(
                            _read_tags_and_cover, file_path, skip_cover=bool(cached_thumb_hash)),
                        timeout=READ_TIMEOUT_SECS)
                except Exception as e:
                    # Not cached. A file that genuinely carries no tags reads
                    # fine and returns an empty dict, which is an answer worth
                    # keeping; this is a drive that did not answer, and writing
                    # "says nothing" for it would make one bad read permanent.
                    log.warning("Tag read failed for %s: %s", file_path, e)
                    tags, duration, cover = {}, None, None
                else:
                    # `cover_seen` is false when the scan was skipped, so a file
                    # whose cover was cached and has since been evicted is
                    # looked at again rather than left without one for good.
                    await self._media_cache.put_audio_meta(
                        entry.id, tags, int(duration) if duration else None,
                        cover_seen=not cached_thumb_hash)

            if duration:
                fields["duration"] = int(duration)

            parsed = title_parse.parse_track_filename(entry.name)
            fields["display_title"] = tags.get("title") or parsed.title or parsed.naive_title
            fields["track_no"] = tags.get("track_no") if "track_no" in tags else parsed.track_no

            artist = tags.get("artist")
            album = tags.get("album")
            if not artist or not album:
                fallback_artist, fallback_album = await asyncio.to_thread(
                    _artist_album_from_ancestors, file_path, root_path)
                artist = artist or fallback_artist
                album = album or fallback_album
            fields["artist"] = artist
            fields["album"] = album

            if cached_thumb_hash:
                fields["thumb_hash"] = cached_thumb_hash
            elif cover:
                thumb_hash = blake3.blake3(cover).hexdigest()
                await self._media_cache.put_thumb(thumb_hash, entry.id, cover)
                fields["thumb_hash"] = thumb_hash

            await on_done(entry.id, fields)