
# Enrich a library

Enrichment adds the information that the volume does not hold: a media
folder's title, plot, ratings, art, and people. The operator asks
metadata providers and writes their answers beside the media. Kodi and
Jellyfin read the resulting `.nfo` files and art files. The volume
remains the source of truth, and the catalog is derived from it.

## 1. Declare a provider

A `MetadataProvider` configures one provider. For providers that require
an API key, store the key in a `Secret` in the same namespace:

    apiVersion: v1
    kind: Secret
    metadata:
      name: tmdb-api-key
      namespace: media
    type: Opaque
    stringData:
      token: <your key>
    ---
    apiVersion: library.liken.sh/v1alpha1
    kind: MetadataProvider
    metadata:
      name: tmdb
      namespace: media
    spec:
      tmdb:
        secretRef:
          name: tmdb-api-key

The providers, and the facts each one serves:

* `tmdb`, The Movie Database: identity, overview, certification, its
  own rating, credits, poster, backdrop, logo, season posters, episode
  stills, trailers, and the people's ids, biographies, and headshots.
  Every library starts here, because identification happens through
  TMDb.
* `omdb`, OMDb: overview, certification, and the IMDb, Rotten
  Tomatoes, and Metacritic ratings. It answers on an IMDb id, which
  TMDb supplies. The free tier allows a thousand calls a day.
* `fanart`, Fanart.tv: art only, and the one source of clear art,
  banners, landscapes, disc art, and season banners.
* `tvmaze`, TVmaze: series only, and it needs no account. Declare it
  with an empty block, `tvmaze: {}`.
* `peertube`, one PeerTube instance: trailers only, and it needs no
  account. Declare it with the address of the instance,
  `peertube: {endpoint: https://tube.example}`. The trailer fact
  searches the instance by title, so pick an instance that publishes
  trailers.
* `archive`, the Internet Archive: trailers only, from its
  `movie_trailers` collection, and it needs no account. Declare it
  with an empty block, `archive: {}`. It holds trailers for many older
  films, and the operator asks it no faster than four times a second.
* `theintrodb`, TheIntroDB: marks only, the intro, recap, credits, and
  preview spans of movies and episodes. It finds a work by its TMDb id,
  and a key is optional. Declare it with an empty block,
  `theintrodb: {}`, or name a `Secret` under `secretRef` to use a key.
  Without a key it answers 500 asks a day for your address. With one it
  answers 1000 a day for your account, and it adds your own pending
  submissions to each answer.
* `introdb`, IntroDB: marks only, the intro, recap, credits, and
  post-credits spans of movies and episodes. It finds a work by its IMDb
  id and needs no account. Declare it with an empty block,
  `introdb: {}`.
* `imdb`, IMDb's published datasets: the IMDb rating of movies, series,
  and episodes, and the principal credits of movies and series, from
  files that IMDb replaces every day. It needs no account and has no
  daily limit. Declare it with an empty block, `imdb: {}`.
  [IMDb ratings and credits from the datasets](#imdb-ratings-and-credits-from-the-datasets)
  describes how the files are read and kept.

The operator checks that each provider answers with one call: when it
starts, when you edit the provider, and then once an hour. A provider
that gives no usable answer, refuses its key, or has no `Secret` or key
is checked every five minutes until it is `Reachable`. The operator reads the
`Secret` for that call, so a key you fix or a `Secret` you create
shows as `Reachable` within five minutes. It also reads the `Secret`
of each `Reachable` provider just before it creates a `Library`'s
`Job`. A `Secret` that you deleted, or that lost its key, then shows as
`NoSecret` at once, and the `Job` leaves that provider out instead of
waiting for a `Secret` that is gone. An OMDb key that has spent
its calls for the day shows as `LimitReached`, not `Refused`, and is
checked once an hour, so the check does not spend calls the key does
not have. A `secretKeyRef` passes the
key to a phase container of the `Library`'s `Job`, so no long-running
pod stores it.

    $ kubectl -n media get metadataproviders
    NAME    PROVIDER   READY   REASON      UPDATED   AGE
    tmdb    tmdb       True    Reachable             3d
    omdb    omdb       False   Refused               3d
    imdb    imdb       True    Reachable   9h        3d

Each change of the `Ready` reason posts one `Event` on the provider,
which `kubectl -n media describe metadataprovider omdb` shows. A check
that gives the same reason again posts none. `NoSecret`, `Refused`,
`LimitReached`, `Unreachable`, `Unavailable`, and the `Cached` reason
`ClaimFailed` post a `Warning`, and the other reasons post a `Normal`
`Event`.

`UPDATED` is for `imdb` alone: the oldest time IMDb replaced one of the
files the provider reads.

`spec.facts` narrows what one account serves.
[MetadataProvider](/docs/reference/metadataproviders/) describes every
field.

## 2. Name the sources on the Library

A `Library` asks the providers it names in `spec.sources`, in that
order:

    spec:
      sources:
        - tmdb
        - omdb
        - fanart

A fact with one value, such as a plot or a certification, takes the
first provider that answers. A fact with a set of values, such as the
genres, takes the union of every provider that answers. The providers
spell some genres differently, so the enricher writes each genre in one
spelling: `Sci-Fi` and `Science-Fiction` become `Science Fiction`,
`Sport` becomes `Sports`, and `Talk` becomes `Talk Show`. TMDb's
combined series genres split in two: `Sci-Fi & Fantasy` becomes
`Science Fiction` and `Fantasy`, `Action & Adventure` becomes `Action`
and `Adventure`, and `War & Politics` becomes `War` and `Politics`. Art takes the
first provider that holds an image. The `Sources` condition on the
`Library` reports whether every name resolves and every fact the
library needs has a provider.

Put `imdb` before `omdb` for the IMDb rating. The datasets rate every
title in one read of one file, and OMDb's thousand calls a day then go
to the plot, the certification, and the Rotten Tomatoes and Metacritic
ratings. Only `imdb` rates episodes, and only when it is the first
source that serves `rating.imdb`:

    spec:
      sources:
        - tmdb
        - imdb
        - omdb

Keep `tmdb` before `imdb` for the credits. IMDb's datasets hold about
nine people for each title, the billed cast and the key crew, and TMDb
holds the full cast, of which the `tmdb` source writes the first 50
people. So `imdb` answers the credits of a title only where
no source before it answered: a title TMDb does not hold, or a
`Library` with no TMDb key.

## 3. What the phases do

Enrichment runs as phases of the `Library`'s `Job`, one container for
each phase. The operator runs one `Job` of a `Library` at a time. A walk
`Job` runs the walk and every phase the `Library`'s sources serve. A
`Job` that fills gaps runs no walk and only the phases whose gaps the
last report counted. The [scanning guide](/docs/guides/scanning/#when-a-scan-runs)
says when each one starts.

The phases are `probe`, which reads each video's streams, `arrival`,
which records when a file was first seen, `identity`, which names each
title, `nfo`, which fills the `.nfo` file, `art`, which downloads the
images, `trailer`, which records where each title's trailers are,
`marks`, which records where each video's intro and credits are,
`contributors`, which fills the people, and, where the `Library` turns
it on, `trailer-files`. A phase runs only where a Ready source of the
`Library` serves one of its facts. `probe` and `arrival` ask no
provider, so they always run. The trickplay and appearances facts are
no phases: each runs in a worker `Job` of its own, which the
[trickplay](#trickplay) and [appearances](#appearances) sections
describe.

All the phases start together, and each one reads its gap again
whenever the rows it reads change. So a phase works on a title as soon
as the phases before it have written that title's rows: `nfo`, `art`,
and `trailer` wait for the id `identity` writes, `marks` for the id and
the length `probe` measures, `contributors` for the people the credits
fact writes, and `trailer-files` for the addresses `trailer` records. The art of the first title lands while
`identity` still works on the rest.

A phase ends when every phase it waits for has ended and one pass after
that found no work. Then it writes a mark on the `Job`'s `phases`
volume, and the `close` container waits for every mark before it writes
the `Job`'s run. Three phases edit the `.nfo` file of a title: `probe`
writes `<fileinfo>`, `identity` writes `<uniqueid>`, and `nfo` writes
its element groups. Each edit takes a lock on the `phases` volume for
that file, so no two of them overwrite each other.

A phase that fails writes a failed mark, and the phases that wait for
it finish what they have and end. The `Job` still succeeds, the
`enrich` run in `status.runs` names the failure, and the failed phase's
titles stay in its gap for the next `Job`.

    kubectl -n media get jobs -l library.liken.sh/library=movies
    kubectl -n media logs job/<job> -c nfo

### Trickplay

The scrub-bar thumbnails are the `trickplay` fact, which runs where
`spec.trickplay.enabled` is set. One title's decode runs for minutes,
so the fact runs in a worker `Job` of its own, outside the `Job` that
walks the `Library`. A walk and a webhook's rescan never wait for a
decode, because the worker runs no catalog agent and holds none of the
`Library`'s catalog claim.

The worker gets its videos from the bus. When every phase of a
`Job` of the `Library` has ended, the `Job`'s close container reads
the trickplay gap from its copy of the catalog, with the `spec.refresh`
time applied, and publishes each video as one retained message on the
broker that media-operator runs, numbered from 0, and then the count
of them:

    liken/library/libraries/<namespace>/<library>/missing/trickplay/<job>/<index>
    liken/library/libraries/<namespace>/<library>/missing/trickplay/<job>/count

Each message holds the video's path, its size, and its length. The
close container waits until the broker holds the count before it
writes the `Job`'s enrich run, and a list the broker did not take is a
failure in that run. One list holds at most 10,000 videos, and the
next `Job` lists the rest. You can read the videos that wait with
`mosquitto_sub`:

    kubectl -n liken-system exec deploy/bus -- mosquitto_sub -v -W 2 \
      -t 'liken/library/libraries/media/movies/missing/#'

When that `Job` has finished, the bus holds its count, and no trickplay
worker of the `Library` runs, the operator starts one, named
`<library>-trickplay-<suffix>`. The worker is an Indexed `Job` with
one completion for each video of the list, so each pod works one
video, and the `Job` controller starts the next index when a pod ends.
Each pod reads the message at its own index, and checks the video
again before it decodes it. It passes over a video that is gone, a
video whose size changed, and a video with an attempt in
`.liken/trickplay.yaml` from after that enrich run finished. When the
video is done, the pod asks the operator to rescan the video's title
folder through the `Library`'s webhook address, so the catalog shows
the tiles within seconds. A pod whose message is gone, because the
broker restarted, has nothing to do and ends. Its video stays in the
gap, and the next `Job` lists it again.

The operator clears a list from the bus when the worker that took it
has finished, and when a newer list of the same fact replaces a list
that no worker took.

The worker has no time limit of its own. A backlog runs to the end of
the list in one `Job`, and a `Job` that runs for 24 hours reaches its
deadline. The videos whose indexes did not run stay in the gap, and the
end of the next `Job` of the `Library` lists them for another worker.
`kubectl get job` shows the worker's completed indexes out of its
count, and `status.failedIndexes` names the indexes that failed. Each
pod has its own log:

    kubectl -n media get jobs -l library.liken.sh/library=movies,library.liken.sh/worker=trickplay
    kubectl -n media get pods -l library.liken.sh/library=movies,library.liken.sh/worker=trickplay

The worker decodes on a GPU when `spec.trickplay.gpuResourceClaimTemplate`
names a `ResourceClaimTemplate`, as the next section describes.
`ffmpeg` decodes through VA-API, and it falls back to software for a
codec the GPU refuses. It does not fall back when the driver cannot
open the GPU at all. With no template, the worker decodes in software.

#### A worker on a GPU

The operator does not build the GPU claim of a worker. A claim can
take any shape that dynamic resource allocation allows, and the choice
of a GPU is the cluster owner's, so the device classes and selectors
stay in your own YAML. You write a `ResourceClaimTemplate` in the
namespace of the `Library`, and the `Library` names it in
`spec.trickplay.gpuResourceClaimTemplate` or
`spec.appearances.gpuResourceClaimTemplate`. Each pod of the worker
claims a device from that template. The field is named for the role
of the claim, the GPU the worker decodes on, so a claim for another
kind of device, such as an NPU for the face models, takes a field of
its own.

Both workers need a GPU whose VA-API driver decodes the library's
video, 10-bit HEVC included, and scales 10-bit frames. A render node
alone does not say what its driver can do. On a fleet of mixed GPUs, a
pod can then land on a GPU that the image's driver cannot open, such as
an AMD GPU when the image holds the Intel driver. `ffmpeg` fails on
every video before it reads a frame, and the worker records a miss for
each one, which lasts 30 days. So the template pairs the render node
with a `media.liken.sh` capability device of the same GPU, and the
scheduler places the pod only on a GPU whose driver states that it
decodes and scales 10-bit video.
[Claim a GPU by what it decodes](https://liken.sh/media/docs/guides/claim-a-gpu-by-capability/)
describes the capability devices and writes the `decode-10bit` class.
The `gpu-render` class here selects every render node that `liken`
publishes:

```yaml
apiVersion: resource.k8s.io/v1
kind: DeviceClass
metadata:
  name: gpu-render
spec:
  selectors:
    - cel:
        expression: |
          device.driver == "liken.sh" &&
          has(device.attributes["liken.sh"].renderNode)
---
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  name: trickplay-gpu
  namespace: media
spec:
  spec:
    devices:
      requests:
        - name: gpu
          exactly:
            deviceClassName: gpu-render
        - name: decodes
          exactly:
            deviceClassName: decode-10bit
      constraints:
        - requests: [gpu, decodes]
          matchAttribute: resource.kubernetes.io/pciBusID
```

```yaml
apiVersion: library.liken.sh/v1alpha1
kind: Library
metadata:
  name: movies
  namespace: media
spec:
  trickplay:
    enabled: true
    gpuResourceClaimTemplate: trickplay-gpu
```

The appearances worker takes a template of the same shape in
`spec.appearances.gpuResourceClaimTemplate`. On a fleet where every
GPU's driver decodes the library's video, the `gpu` request alone is
enough.

This template needs a `liken` release that publishes
`resource.kubernetes.io/pciBusID` on its render nodes. On an older
release, the claim never allocates, and the worker's pods stay
`Pending`.

The operator reads only whether the template exists. When the
`Library` names a template that its namespace does not hold, the
operator starts no worker of that fact, because the worker's pods
would stay `Pending` until the `Job`'s 24-hour deadline. The
`GPUClaimTemplates` condition of the `Library` is then `False` with
the reason `ClaimTemplateNotFound`, and its message names the
template:

    kubectl -n media get library movies -o jsonpath='{.status.conditions[?(@.type=="GPUClaimTemplates")]}'

The operator watches the templates, so the next worker starts as soon
as you create the template. A template that exists but that no node
can allocate leaves the worker's pods `Pending`, and their events say
so.

The `Library` schema has no `render` field. An earlier release of the
operator built a template for each worker from that field, named
`<library>-trickplay` or `<library>-appearances`, with the `Library`
as its owner. The operator deletes each such template the first time
it reconciles the `Library`, and leaves a template you wrote at the
same name.

#### A worker on several nodes

At the default parallelism, the worker decodes one video at a time. A
large library can decode several videos at once, on several nodes'
GPUs, with `spec.trickplay.parallelism` or
`spec.appearances.parallelism`, the number of pods the worker `Job`
runs at once, from 1 to 16:

```yaml
spec:
  appearances:
    enabled: true
    parallelism: 3
    gpuResourceClaimTemplate: appearances-gpu
```

Each pod claims its own device from the template that
`gpuResourceClaimTemplate` names. The pods prefer different nodes.
When fewer nodes offer the device than the worker runs pods, two pods
share a node and its device, because `liken` publishes a render node
for many claims at once. Every pod goes through the scheduler on its
own, so a node's taint and the pod's resource requests apply to each
one. Kubernetes retries a failed pod up to twice on its own video, and
leaves the other indexes alone. The 24-hour deadline is for the whole
`Job`. The worker counts as running until every index has ended, so
the next worker waits for the last pod.

Two pods can work two episodes of one season at once, and each writes
the season folder's ledger. Each write reads the ledger again just
before it replaces the file, and applies the pod's own entry to what
the other pod left, so both attempts stay. The two writes can still
collide in the moment between that read and the rename. Each pod asks
for a rescan of the series folder, and the operator holds the requests
until its next pass, which walks the folder once.

The tile directory is the one Jellyfin reads and writes. So the worker
accepts a directory Jellyfin made first and leaves it alone. The one
exception is a directory older than the file beside it: when a new
file replaced the video at the same path, tiles made before it arrived
place each thumbnail at the earlier file's times. The worker decodes
the new file and replaces the whole directory. The
[scanning guide](/docs/guides/scanning/#a-file-replaced-at-the-same-path)
says how the walk finds a replaced file. To make the tiles of any
title again, delete its `.trickplay` directory, and the next walk opens
its gap, and the end of that `Job` starts a worker.

### Appearances

The `appearances` fact records which credited actor is on screen at
each second of a feature. It runs where `spec.appearances.enabled` is
set. The fact asks no provider: it matches the faces in the video with
the headshots that the people facts put in `.contributors/`. A first
pass decodes every frame of every feature, so the fact runs in a
worker `Job` of its own, named `<library>-appearances-<suffix>`, in the
same way as the [trickplay](#trickplay) worker. The `Job` that walks
the `Library` publishes the appearances gap on the bus under
`missing/appearances`, and each pod of the worker reads one video,
checks it again before it opens it, and asks the operator to rescan
its title folder when it is done.
`spec.appearances.parallelism` runs the worker on several nodes at
once, as [a worker on several nodes](#a-worker-on-several-nodes)
describes.

A feature is in the gap when the probe gave it a length and its title
credits at least one actor whose entry holds a headshot. The credits
fact credits a series and not each episode, so an episode takes the
cast of its series. No source records an episode's guest stars yet, so
a guest star has no headshot to match. A title whose cast has no
headshot waits until the headshot fact writes one.

For each video, the worker runs two passes of the `appearances` tool:

1. `appearances detect` decodes every frame of the video with `ffmpeg`,
   finds the faces with the YuNet detector in every keyframe and in a
   frame every second, embeds each face with the SFace model, and
   writes the detections record `.liken/appearances/<file>.jsonl`
   beside the video. This is the only pass that opens the video. The
   detector keeps each face it scores at 0.8 or more, and the record
   marks each keyframe. The record's first line names the record's
   format, the file's size, and the hash of each model. When a record
   on the volume names the current format,
   `liken.sh/appearances/detections/v2`, the file's size, and the models
   of the image, the worker uses it and does not decode again. A record
   of any other format holds the keyframes alone, so the worker decodes
   its video again.
2. `appearances match` embeds the headshot of each actor that the
   title folder's `.liken/credits.yaml` names, and names the faces in
   two passes. The headshot pass names a face for the closest actor when
   the similarity is at least 0.363, the threshold OpenCV publishes for
   SFace, and leads the next closest actor by at least 0.05. The film
   pass then names a face the headshots missed, such as a face in
   profile or decades younger than the headshot, when at least 2 faces
   the headshots named in other samples are close to it, all of them
   name one actor, and they are at least 30% of the faces close to it.
   The match takes 1 to 5 seconds, and it runs on the CPU.

The worker writes the answer to `.liken/appearances.yaml` beside the
video, one entry per file:

    appearances:
        - path: Example Movie (2019).mkv
          size: 4831838208
          embedder: {name: face_recognition_sface_2021dec, sha256: 0ba9fbfa...}
          threshold: 0.363
          margin: 0.05
          gallery:
            - {contributor: .contributors/ad/ada-quill, name: Ada Quill, headshot: found, sha256: 5d2c...}
            - {contributor: .contributors/bo/bo-reyes, name: Bo Reyes, headshot: missing}
          unmatched:
            - {contributor: .contributors/bo/bo-reyes, name: Bo Reyes, headshot: missing}
          named: {headshot: 4326, film: 921}
          people: 21
    attempts:
        - path: Example Movie (2019).mkv
          at: 2026-10-01T12:00:00Z
          result: found

The entry records the inputs of the answer: the size of the file, the
embedding model, the threshold, the margin, and each actor's headshot
with its hash. `named` counts the faces each pass named, and `people`
counts the actors they name. `unmatched` lists the actors who cannot be
named, because their entry holds no headshot or the detector found no
face in it.

The faces themselves are in `.liken/appearances/<file>.matches.json`
beside the video, one JSON document of 0.4 to 0.8 MB for a film. Each
face is one observation: the time of its sample in seconds, the face's
place in that sample's line of the detections record, the actor, the
pass that named the face, the similarity to the actor's headshot, and
the closest other actor's similarity. A face named `by` `film` can be
under the threshold. The ledger holds only the counts, because a walk
reads every ledger of a folder on each pass, and the faces of one film
would make the ledger entry about 700 KB. A failed pass is an `error`
attempt whose `reason` holds the tool's own error text, and the worker
tries the file again after a day.

The match also writes `.liken/appearances/<file>.spans.json` beside
the video, for the screens. It holds spans, not samples: each run of
samples that names the same actors, with the actors left to right in
the order they stand in the picture, and each actor's name, character,
headshot, and dates. Each sample stands for the time halfway to its
neighbours, and never past the next keyframe, where a cut can fall. The
actors named last hold through samples that name nobody for up to 10
seconds, and two samples in a row with no face end the span. When a person plays a video whose answer is
`found`, the screen's browser names the spans file and the library's
`.contributors/` in the `Play`. The display then shows who is on screen
when the film pauses. The
[tool's README](https://github.com/liken-sh/liken/tree/main/library-operator/appearances#the-spans-file)
gives the file's shape.

A found answer stands until a new file replaces the video at its path.
The answer does not change when a headshot or the credits change. To
match a title again, delete its `.liken/appearances.yaml`. The next
walk opens the gap, and the worker uses the detections records on the
volume and runs only the match. To match every title again, set
`spec.refresh.appearances` to the current time, as the
[When a fact asks again](#when-a-fact-asks-again) section describes
for every fact. A refresh also decodes again each video whose record
is of an earlier format, and the match then rewrites its ledger entry
and its spans file in place.

    kubectl -n media get jobs -l library.liken.sh/library=movies,library.liken.sh/worker=appearances

The worker runs on the `library-operator-appearances` image, which
carries the tool, Intel's OpenVINO runtime, Intel's OpenCL runtime for
the GPU, and the two models from
[OpenCV Zoo](https://github.com/opencv/opencv_zoo): YuNet under the MIT
license and SFace under the Apache 2.0 license. The operator names the
image at its own tag, and `APPEARANCES_IMAGE` on the operator's
`Deployment` names another. The container requests one core of CPU and
may take `1536Mi` of memory. The largest measured run held 740 MB in
the tool and `ffmpeg` together, for a 4K file decoded in software. On a
laptop, the detect pass took 1.5 to 4 minutes for a 1080p film and 6
minutes for a 4K film through VA-API, and 34 minutes for the 4K film in
software. The worker allows each detect pass 6 hours.

With no template, the worker decodes and runs the models on the CPU.
`spec.appearances.gpuResourceClaimTemplate` names a
`ResourceClaimTemplate`, as [A worker on a GPU](#a-worker-on-a-gpu)
describes for both workers. With the claim, `ffmpeg` decodes and scales the video on
the render node through VA-API, and OpenVINO runs the models on the
Intel GPU. Some GPUs decode a format that their video processor cannot
scale, such as 10-bit HEVC on an older Intel GPU. The worker tries the
GPU's scale on the first 2 seconds of each video, and where that fails,
the GPU decodes and the CPU scales. A file the render node refuses to
decode is decoded again in software. OpenVINO compiles the models for the GPU when the worker
starts, which took seconds in the measurements, and it keeps the
compiled kernels in an `emptyDir` that the detect and match passes of
the pod's one video share.

#### Checking the faces by hand

`appearances review` shows a person what the match named, on a copy of
a title folder on a workstation. It runs no model and writes nothing
into the folder. Never run the tool on the library itself. The copy
needs the title folder with its `.liken` directory, and the
`.contributors/` entries its credits name, at the same paths under a
common root. The
[tool's README](https://github.com/liken-sh/liken/tree/main/library-operator/appearances)
says how to build it and how to install the runtime.

    appearances match "Movies/Example Movie (2019)"
    appearances review "Movies/Example Movie (2019)" --sheets 24 --play

`--play` opens `mpv` with a chapter for each span and a box around each
face: green for a named face, yellow for a face near the threshold or
the margin, and red for a face no actor is close to. `--sheets 24`
writes contact sheets of the faces named for each actor, weakest
first. `--threshold` and `--margin` show the result of other values
with no pass over the video. For an episode, run the match on the
season folder and name the series' credits:

    appearances match "Series/Example Show/Season 01" --credits "Series/Example Show/.liken/credits.yaml"

### Identification

The identity fact asks TMDb for the folder's title and runs a fixed
sequence of tests: the title, then the year, then a year on either side,
then, for a series, the episode names, then the runtime within five
minutes. The episode test reads the episode titles from the file names
of the folder's first season. It keeps a candidate whose season on TMDb
has at least two of them and at least half. So a series folder named
with the title alone identifies without an `.nfo` file. One survivor is the
answer, and its reason is recorded. Several survivors become candidates
in `.liken/identity.yaml`, and the title counts in `status.waiting`
until a person names the right `uniqueid` in the `.nfo`. A title no
provider can name counts in `status.unresolved`.

After identifying a title through TMDb, the identity phase requests its IMDb
and TVDB ids and writes any returned ids into the `.nfo` file. OMDb uses
the IMDb id to look up the title. Fanart.tv uses the TMDb id for a
movie and the TVDB id for a series. These ids identify the title;
OMDb and Fanart.tv still require their own API keys, configured through
each provider's `secretRef`.

### The write rule

Each fact owns a fixed group of elements in the `.nfo` file and writes
nothing outside it. The `overview` fact owns the plot, the tagline,
the genres, the studios, the premiere date, and the runtime. Each
rating owns its one element. `credits` owns the actors, directors,
and writers. Before a fact writes again, it hashes the group as it is
now and compares that hash with the one it recorded after its last
write. If the hashes differ, another writer changed that group. The
fact records a fight and does not write the group. `status.fights`
counts the fights.

An art file that already exists is never replaced. The fact records it
as answered and downloads nothing. The one exception is an episode's
thumbnail that is older than the episode's file, where a new file
replaced the video at the same path. That thumbnail is a frame of the
earlier file, so the fact downloads the still again and writes it over
the old one.

Every file lands whole: the writer fills a temporary beside it and
renames the temporary into place. Two clusters can mount one library,
each with its own `Library` over the same root, and their writers can
reach one file at once. A writer that changes a file it read, an
`.nfo` file, a `.liken` ledger, or a person's `contributor.yaml`,
reads the file again just before its rename. Where another writer
changed the file in between, it applies its change to what that
writer left, so neither change is lost. The two clusters can still do
the same work twice. The check keeps each other's writes, and it does
not divide the work between them.

### People

The `credits` fact writes each title's cast and crew into
`.liken/credits.yaml`, and it gives each credited person one entry in
`.contributors/` at the library root. A credit looks for its entry by
id first. When the catalog holds an entry with one of the credit's
ids, the credit names that entry, whatever the spelling of the name.
A credit that holds an IMDb id and no TMDb id asks TMDb for the TMDb
id first, when the `Library`'s sources name a `Ready` `tmdb` provider.
When no id finds an entry, the credit uses the entry at the slug of
the name. A second person of the same name gets the slug with an id,
such as `nora-vance-tmdb-992`.

A credit from IMDb's datasets names a person by the IMDb id alone. So
it finds the entry a TMDb credit wrote for the same person by that id,
or through the TMDb id that TMDb's find call gives for it, and the
person keeps one entry. An entry the datasets created holds the IMDb id,
and `contributor.ids` fills its TMDb id, its biography, and its
headshot the same way.

The `contributors` container fills each entry. `contributor.ids`
writes the birth date, the death date, and the person's ids in other
databases into `contributor.yaml`. For an entry that holds only an
IMDb id, it asks TMDb for the TMDb id first. `contributor.biography`
and `contributor.headshot` write `biography.txt` and `headshot.jpg`
beside the entry where no file of that name exists.

Two entries can hold one id, for example when two providers spell one
name two ways. After it writes the ids, `contributor.ids` merges each
group of entries that share an id into one entry:

- The entry at the slug of its own name, with no id suffix, stays. If
  no entry of the group is at such a slug, the entry with the most ids
  stays.
- The entry that stays gets every id of the group. `biography.txt` and
  `headshot.jpg` move to it where it has no file of that name.
- Each other entry keeps a `contributor.yaml` with one field,
  `mergedInto`, the path of the entry that stays. The catalog shows no
  person for such an entry.
- The next `Job` of the `Library` moves each credit that names a
  removed entry to the entry that stays. The nfo phase moves the
  credits, and it can end before the merge, so the move waits for that
  `Job`. The contributors phase of the same `Job` then deletes each
  removed entry that no credit names.

The merge leaves a group whole in two cases, and
`.liken/contributor.merge.yaml` in each entry of the group records
the reason. A `held` attempt means a person edited a
`contributor.yaml` of the group, and `status.fights` counts each entry
of it. A `conflict` attempt means two entries hold two different ids
in one scheme, so they are two people and one of them holds a wrong
id. The merge asks about the group again after thirty days. To hand
an edited entry back to the merge sooner, delete
`.liken/contributor.ids.yaml` and `.liken/contributor.merge.yaml` in
that entry. The next walk opens the gap, and `contributor.ids` then
reads the file as its own.

### Trailers

The `trailer` fact records links, not files. It asks every source
that serves it and writes what it finds to `.liken/trailer.yaml`
beside the title, one entry per video: the provider, the provider's
own key, the site the video plays from, the page to watch it on, its
name, its kind (`trailer`, `teaser`, or `spot`), its language, its
resolution where the provider states one, and a score from 1 to 100
with a one-line reason. The catalog's `trailers` table holds the same
rows.

The score ranks the videos of one title, so a screen can take the first.
A video keyed by the title's own id, as TMDb's are, starts at 100. A
video found by a search starts at 90 when its name has the same title
and year. It starts at 60 when the name has the title and no year at
all. A search result whose name has another title or another year is not
recorded. Then the kind takes points off, a teaser less than a TV spot.
A language the household does not prefer takes points off, and so does a
TMDb video that is not marked official. The reason says which of these
applied. Clips, featurettes, and other extras are not recorded at all.
Within one provider, videos with the same name collapse to the best one,
and at most five videos per provider are kept for a title. An Internet
Archive item is kept only when its own video runs eight minutes or less,
because the `movie_trailers` collection holds whole films beside the
trailers.

The preferred languages are the library's `spec.languages`, then the
household's `audioLanguages` from the media operator's
`MediaPreferences`, and `en` when neither names any.

A trailer file beside the title still plays as before.

### Intro and credits marks

The `marks` fact records where a video file's intro, recap, credits,
and preview are. It asks every source that serves it about each main
video of an identified movie, and of each episode of an identified
series, once the `probe` fact has measured the file's length. It writes
what it finds to `.liken/marks.yaml` in the folder that holds the file,
the folder whose `.liken/probe.yaml` records the same file:

    marks:
      - path: A Series - S01E02.mkv
        kind: intro
        end: 107000
        source: theintrodb
      - path: A Series - S01E02.mkv
        kind: intro
        start: 7007
        end: 106482
        source: theintrodb
      - path: A Series - S01E02.mkv
        kind: credits
        start: 3253000
        end: 3316000
        source: theintrodb

Each entry is one span: the file, the kind, the start and the end in
milliseconds from the start of the file, and the provider that answered.
An absent `start` is the start of the file, and an absent `end` is the
end of the file. A provider can answer several candidates for one kind,
from several submissions or several releases of the work, and the fact
records every one exactly as the provider answered it. It chooses none.
The catalog's `marks` table holds one row per span, and the media
browser sends every span to the `Play`, where the player in
`media-operator` reads them and offers the skip.

TheIntroDB reads the file's length from the ask, and it answers the
spans of the release whose length is closest, such as the theatrical
cut or the extended one. IntroDB reads no length. A file that holds two
episodes is asked about neither, because each provider places an
episode's spans in that episode's own file.

The community adds marks for a new episode or film in the days after
it comes out, so the fact asks again about a new work sooner than
about an old one. The release date the catalog holds for the movie or
the episode sets the wait after a find or a miss:

| Released | Asked again after |
|---|---|
| in the last 7 days | 1 day |
| 8 to 90 days ago | 7 days |
| more than 90 days ago, or no date | 30 days |

A file that holds two episodes takes the later of their dates. An
error waits one day, whatever the date.

The table applies only where every provider the `Library` names
answered. Where one provider failed, or was not asked, the attempt
records the result `partial`. The file keeps the spans the other
providers answered, keeps the spans the silent provider answered
before, and waits one day, so that provider is asked again the next
day.

A provider that answers `429` waits for the reset its headers name and
asks again. When a provider has spent its allowance for the day, the
reset is hours away, and the fact asks that provider nothing more in
this run. Every later file of the run is `partial`, and the run stops
when every provider has spent its allowance. A file that every provider
answered keeps the window in the table, so the next day's allowance
goes to the files the spent provider missed and not again to the ones
it answered.

Neither Jellyfin nor Kodi reads these marks. Jellyfin keeps its media
segments in its own database, and Kodi's `.edl` file is a different
format, so the fact writes no file other than its ledger.

### Trailer files

The `trailerfile` fact pulls one video file per title. It is off unless
`spec.trailers.enabled` is `true`, and it pulls nothing until the
`trailer` fact has recorded links.

The file lands at `<title>/trailers/<name>.mp4`. The name is the
trailer's own name with every character a file name cannot hold taken
out, and its length is capped.

The fact takes the highest-scored trailer whose site the operator can
fetch from. Today those sites are the Internet Archive and PeerTube,
never YouTube. From that trailer's files it takes the tallest file that
is no taller than the title's own feature, or the shortest above it
where none fits.

The fact never pulls for a title that already holds a trailer file
anywhere under its folder. A trailer a person placed by hand stays, and
the fact records nothing.

Every pull is remuxed to MP4 and checked with `ffprobe` before it
lands. The check requires a video stream and a length between 10
seconds and 8 minutes, so a TV spot passes and a whole film does not.
A file that fails the check never reaches a name the walk reads, and
the attempt records the error.

**A trailer is tens of megabytes per title, so a library of any size
adds gigabytes to the volume the first time this fact runs.**

The `trailer-files` phase of the `Library`'s `Job` runs the fact, and
two titles pull at once inside it. The phase starts no new title
fifteen minutes into a run, and the next `Job` goes on with the rest. The fact takes no `spec.refresh`.

### IMDb ratings and credits from the datasets

IMDb publishes its datasets as gzipped files at `datasets.imdbws.com`
for personal and non-commercial use. No `liken` image carries them, so
each cluster downloads them from IMDb. IMDb's terms require this credit
where the data is shown:

> Information courtesy of IMDb (https://www.imdb.com). Used with
> permission.

The `nfo` container reads its whole `rating.imdb` gap first. Then it
reads `title.ratings` once, from start to end, and keeps only the rows
of the titles in the gap, so one read serves one title or ten thousand.
The read starts when the container starts, and the other facts run
while it works. A run with no `rating.imdb` gap sends no request.

A movie or a series is found by the IMDb id in its `.nfo` file. An
episode is found by its own IMDb id where its `.nfo` file names one.
Otherwise the container reads `title.episode` once and finds the
episode by its season and episode numbers under the series' IMDb id. It
records the id it found in `.liken/rating.imdb.yaml`, so a later run
does not read `title.episode` again for that episode. A file that holds
two episodes gets no rating, because its ledger entry names one episode.

The credits read two files in sequence. The container reads
`title.principals` once and keeps the rows of the titles in its
`credits` gap, then reads `name.basics` once for the names of the
people those rows name. It does not read `name.basics` when every one
of those people already has an entry in `.contributors/`, because the
entry holds the name. Decompressing and reading the two files took
about 32 seconds on a cloud host in September 2026, however many titles
they fill, and the downloads add to that on the first run of a node. `actor`, `actress`, and `self`
rows are the cast, in IMDb's order, with the characters as the role.
`director` and `writer` rows are the crew. The other categories, such
as `producer` and `composer`, have no part in the credits and are not
written, and neither are `archive_footage` and `archive_sound`. The
credits are not asked for again on a timer. Set `credits` in
`spec.refresh` to ask again.

The rating is written into the `.nfo` file as OMDb writes it: the
`imdb` rating, out of 10, with the vote count. The container writes the
file only when the rating at one decimal changed. A change in the vote
count alone writes nothing, so a library's `.nfo` files do not all
change every month. The attempt records the `Last-Modified` time of the
`title.ratings` copy it read.

A rating from the datasets is asked for again after 30 days, but only
when IMDb has published a newer `title.ratings` than the one the
attempt read. A rating OMDb wrote counts as older, so a library that
moves from `omdb` to `imdb` moves each title within 30 days. A rating
that another tool, such as Radarr or Jellyfin, wrote into the `.nfo`
file has no attempt. The next `nfo` container that reads the datasets
reads that rating once and records an attempt, and it writes the file
only when the rating at one decimal differs. The gap count in the
status does not include these ratings, so they start no `Job` of their
own. While the
provider's `Stale` condition is `True`, IMDb has published no newer
file, and the operator starts no `Job` that fills gaps for these ratings
alone.

With a per-node `StorageClass` in the cluster, the operator keeps the
files on the claim `<provider>-datasets`, which the `nfo` container
mounts. Each node's copy fills the first time a run on that node needs
a file. Every later run asks IMDb with the copy's `ETag`, gets `304 Not
Modified` while the file is unchanged, and reads the copy from disk. A
new version replaces the copy only after the whole file arrived and its
gzip checksum is correct. Two runs on one node download a file once:
the second waits for the first and reads its copy. When the claim is
full or cannot be written, the container reads the file from IMDb and
logs the filesystem's error. With no per-node class, every run reads
the files from IMDb.

    kubectl -n media get metadataprovider imdb -o jsonpath='{.status.imdb.datasets}'
    kubectl -n media logs job/<job> -c nfo | grep -E 'title.ratings|title.principals'

### When a fact asks again

A miss lasts for thirty days and an error for one day, then the fact
asks again. The `marks` fact asks about a new work sooner, as the
section above describes. An attempt made before a title's release date
lasts only until that date. An attempt also stops counting when the walk finds that a new file
replaced the video at its path, or that the file or directory the
attempt wrote is gone. The
[scanning guide](/docs/guides/scanning/#a-file-replaced-at-the-same-path)
describes both. To ask one fact again for every title now, set its
time in `spec.refresh`:

    spec:
      refresh:
        overview: "2026-09-06T00:00:00Z"

Every attempt of that fact before the time no longer counts. A refresh
starts a `Job` that fills gaps, without a walk, as soon as the
`Library` has no other `Job` running, and a refresh set while a `Job`
runs starts another one after it. The fact
rewrites its own files and rows in place, and nothing is deleted.
The `appearances` key works the same way: the reopened videos are in
the gap the worker reads after the `Job` that fills gaps, and the
worker matches them again. The worker takes the videos whose detections
record is on the volume first, because each of them needs only the
match.

`kubectl liken library reenrich movies` writes that field for you and
asks every fact again. Add `--only overview` to reopen one fact, and
`-n` to select the namespace. The command reads your kubeconfig, and
in bash it completes the library names.

The same map takes one key that is not a fact: `scan` asks for a full
walk of the library, and `kubectl liken library rescan movies` writes
it. The [scanning guide](/docs/guides/scanning/) describes what it does.

## 4. The Jellyfin handover

Jellyfin reads the `.nfo` files and the art that this operator writes, under the
same names its own scraper uses. **Turn off Jellyfin's
`SaveLocalMetadata` for a library this operator enriches.** With it
on, two writers change the same `.nfo` files, and the fight check leaves
the file to Jellyfin.

## Reading progress

    kubectl -n media get library movies -o jsonpath='{.status.gaps}'
    kubectl -n media get library movies -o jsonpath='{.status.waiting} {.status.unresolved} {.status.fights}'
    kubectl -n media get library movies -o jsonpath='{.status.conditions[?(@.type=="Sources")]}'

`status.gaps` counts, per fact, the rows still to fill. Between walks,
the operator starts a `Job` that fills gaps with the phases whose counts
are above zero, once a refresh time or a source provider that turned
`Ready` has given them work since the last `Job` started. `trailer-files`
runs in such a `Job` while its count is above zero, because each run
stops at its time limit. The `trickplay` and `appearances` counts
start no `Job` that fills gaps: each starts its fact's worker when a
`Job` of the `Library` ends.

