Enrich a library
Enrichment adds the information that the volume does not hold: a media
folder’s title, plot, ratings, art, and people. The operator asks
metadata providers and writes their answers beside the media. Kodi and
Jellyfin read the resulting .nfo files and art files. The volume
remains the source of truth, and the catalog is derived from it.
1. Declare a provider
A MetadataProvider configures one provider. For providers that require
an API key, store the key in a Secret in the same namespace:
apiVersion: v1
kind: Secret
metadata:
name: tmdb-api-key
namespace: media
type: Opaque
stringData:
token: <your key>
---
apiVersion: library.liken.sh/v1alpha1
kind: MetadataProvider
metadata:
name: tmdb
namespace: media
spec:
tmdb:
secretRef:
name: tmdb-api-key
The providers, and the facts each one serves:
tmdb, The Movie Database: identity, overview, certification, its own rating, credits, poster, backdrop, logo, season posters, episode stills, trailers, and the people’s ids, biographies, and headshots. Every library starts here, because identification happens through TMDb.omdb, OMDb: overview, certification, and the IMDb, Rotten Tomatoes, and Metacritic ratings. It answers on an IMDb id, which TMDb supplies. The free tier allows a thousand calls a day.fanart, Fanart.tv: art only, and the one source of clear art, banners, landscapes, disc art, and season banners.tvmaze, TVmaze: series only, and it needs no account. Declare it with an empty block,tvmaze: {}.peertube, one PeerTube instance: trailers only, and it needs no account. Declare it with the address of the instance,peertube: {endpoint: https://tube.example}. The trailer fact searches the instance by title, so pick an instance that publishes trailers.archive, the Internet Archive: trailers only, from itsmovie_trailerscollection, and it needs no account. Declare it with an empty block,archive: {}. It holds trailers for many older films, and the operator asks it no faster than four times a second.theintrodb, TheIntroDB: marks only, the intro, recap, credits, and preview spans of movies and episodes. It finds a work by its TMDb id, and a key is optional. Declare it with an empty block,theintrodb: {}, or name aSecretundersecretRefto use a key. Without a key it answers 500 asks a day for your address. With one it answers 1000 a day for your account, and it adds your own pending submissions to each answer.introdb, IntroDB: marks only, the intro, recap, credits, and post-credits spans of movies and episodes. It finds a work by its IMDb id and needs no account. Declare it with an empty block,introdb: {}.imdb, IMDb’s published datasets: the IMDb rating of movies, series, and episodes, and the principal credits of movies and series, from files that IMDb replaces every day. It needs no account and has no daily limit. Declare it with an empty block,imdb: {}. IMDb ratings and credits from the datasets describes how the files are read and kept.
The operator checks that each provider answers with one call: when it
starts, when you edit the provider, and then once an hour. A provider
that gives no usable answer, refuses its key, or has no Secret or key
is checked every five minutes until it is Reachable. The operator reads the
Secret for that call, so a key you fix or a Secret you create
shows as Reachable within five minutes. It also reads the Secret
of each Reachable provider just before it creates a Library’s
Job. A Secret that you deleted, or that lost its key, then shows as
NoSecret at once, and the Job leaves that provider out instead of
waiting for a Secret that is gone. An OMDb key that has spent
its calls for the day shows as LimitReached, not Refused, and is
checked once an hour, so the check does not spend calls the key does
not have. A secretKeyRef passes the
key to a phase container of the Library’s Job, so no long-running
pod stores it.
$ kubectl -n media get metadataproviders
NAME PROVIDER READY REASON UPDATED AGE
tmdb tmdb True Reachable 3d
omdb omdb False Refused 3d
imdb imdb True Reachable 9h 3d
Each change of the Ready reason posts one Event on the provider,
which kubectl -n media describe metadataprovider omdb shows. A check
that gives the same reason again posts none. NoSecret, Refused,
LimitReached, Unreachable, Unavailable, and the Cached reason
ClaimFailed post a Warning, and the other reasons post a Normal
Event.
UPDATED is for imdb alone: the oldest time IMDb replaced one of the
files the provider reads.
spec.facts narrows what one account serves.
MetadataProvider
describes every
field.
2. Name the sources on the Library
A Library asks the providers it names in spec.sources, in that
order:
spec:
sources:
- tmdb
- omdb
- fanart
A fact with one value, such as a plot or a certification, takes the
first provider that answers. A fact with a set of values, such as the
genres, takes the union of every provider that answers. The providers
spell some genres differently, so the enricher writes each genre in one
spelling: Sci-Fi and Science-Fiction become Science Fiction,
Sport becomes Sports, and Talk becomes Talk Show. TMDb’s
combined series genres split in two: Sci-Fi & Fantasy becomes
Science Fiction and Fantasy, Action & Adventure becomes Action
and Adventure, and War & Politics becomes War and Politics. Art takes the
first provider that holds an image. The Sources condition on the
Library reports whether every name resolves and every fact the
library needs has a provider.
Put imdb before omdb for the IMDb rating. The datasets rate every
title in one read of one file, and OMDb’s thousand calls a day then go
to the plot, the certification, and the Rotten Tomatoes and Metacritic
ratings. Only imdb rates episodes, and only when it is the first
source that serves rating.imdb:
spec:
sources:
- tmdb
- imdb
- omdb
Keep tmdb before imdb for the credits. IMDb’s datasets hold about
nine people for each title, the billed cast and the key crew, and TMDb
holds the full cast, of which the tmdb source writes the first 50
people. So imdb answers the credits of a title only where
no source before it answered: a title TMDb does not hold, or a
Library with no TMDb key.
3. What the phases do
Enrichment runs as phases of the Library’s Job, one container for
each phase. The operator runs one Job of a Library at a time. A walk
Job runs the walk and every phase the Library’s sources serve. A
Job that fills gaps runs no walk and only the phases whose gaps the
last report counted. The scanning guide
says when each one starts.
The phases are probe, which reads each video’s streams, arrival,
which records when a file was first seen, identity, which names each
title, nfo, which fills the .nfo file, art, which downloads the
images, trailer, which records where each title’s trailers are,
marks, which records where each video’s intro and credits are,
contributors, which fills the people, and, where the Library turns
it on, trailer-files. A phase runs only where a Ready source of the
Library serves one of its facts. probe and arrival ask no
provider, so they always run. The trickplay and appearances facts are
no phases: each runs in a worker Job of its own, which the
trickplay
and appearances
sections
describe.
All the phases start together, and each one reads its gap again
whenever the rows it reads change. So a phase works on a title as soon
as the phases before it have written that title’s rows: nfo, art,
and trailer wait for the id identity writes, marks for the id and
the length probe measures, contributors for the people the credits
fact writes, and trailer-files for the addresses trailer records. The art of the first title lands while
identity still works on the rest.
A phase ends when every phase it waits for has ended and one pass after
that found no work. Then it writes a mark on the Job’s phases
volume, and the close container waits for every mark before it writes
the Job’s run. Three phases edit the .nfo file of a title: probe
writes <fileinfo>, identity writes <uniqueid>, and nfo writes
its element groups. Each edit takes a lock on the phases volume for
that file, so no two of them overwrite each other.
A phase that fails writes a failed mark, and the phases that wait for
it finish what they have and end. The Job still succeeds, the
enrich run in status.runs names the failure, and the failed phase’s
titles stay in its gap for the next Job.
kubectl -n media get jobs -l library.liken.sh/library=movies
kubectl -n media logs job/<job> -c nfo
Trickplay
The scrub-bar thumbnails are the trickplay fact, which runs where
spec.trickplay.enabled is set. One title’s decode runs for minutes,
so the fact runs in a worker Job of its own, outside the Job that
walks the Library. A walk and a webhook’s rescan never wait for a
decode, because the worker runs no catalog agent and holds none of the
Library’s catalog claim.
The worker gets its videos from the bus. When every phase of a
Job of the Library has ended, the Job’s close container reads
the trickplay gap from its copy of the catalog, with the spec.refresh
time applied, and publishes each video as one retained message on the
broker that media-operator runs, numbered from 0, and then the count
of them:
liken/library/libraries/<namespace>/<library>/missing/trickplay/<job>/<index>
liken/library/libraries/<namespace>/<library>/missing/trickplay/<job>/count
Each message holds the video’s path, its size, and its length. The
close container waits until the broker holds the count before it
writes the Job’s enrich run, and a list the broker did not take is a
failure in that run. One list holds at most 10,000 videos, and the
next Job lists the rest. You can read the videos that wait with
mosquitto_sub:
kubectl -n liken-system exec deploy/bus -- mosquitto_sub -v -W 2 \
-t 'liken/library/libraries/media/movies/missing/#'
When that Job has finished, the bus holds its count, and no trickplay
worker of the Library runs, the operator starts one, named
<library>-trickplay-<suffix>. The worker is an Indexed Job with
one completion for each video of the list, so each pod works one
video, and the Job controller starts the next index when a pod ends.
Each pod reads the message at its own index, and checks the video
again before it decodes it. It passes over a video that is gone, a
video whose size changed, and a video with an attempt in
.liken/trickplay.yaml from after that enrich run finished. When the
video is done, the pod asks the operator to rescan the video’s title
folder through the Library’s webhook address, so the catalog shows
the tiles within seconds. A pod whose message is gone, because the
broker restarted, has nothing to do and ends. Its video stays in the
gap, and the next Job lists it again.
The operator clears a list from the bus when the worker that took it has finished, and when a newer list of the same fact replaces a list that no worker took.
The worker has no time limit of its own. A backlog runs to the end of
the list in one Job, and a Job that runs for 24 hours reaches its
deadline. The videos whose indexes did not run stay in the gap, and the
end of the next Job of the Library lists them for another worker.
kubectl get job shows the worker’s completed indexes out of its
count, and status.failedIndexes names the indexes that failed. Each
pod has its own log:
kubectl -n media get jobs -l library.liken.sh/library=movies,library.liken.sh/worker=trickplay
kubectl -n media get pods -l library.liken.sh/library=movies,library.liken.sh/worker=trickplay
The worker decodes on a GPU when spec.trickplay.gpuResourceClaimTemplate
names a ResourceClaimTemplate, as the next section describes.
ffmpeg decodes through VA-API, and it falls back to software for a
codec the GPU refuses. It does not fall back when the driver cannot
open the GPU at all. With no template, the worker decodes in software.
A worker on a GPU
The operator does not build the GPU claim of a worker. A claim can
take any shape that dynamic resource allocation allows, and the choice
of a GPU is the cluster owner’s, so the device classes and selectors
stay in your own YAML. You write a ResourceClaimTemplate in the
namespace of the Library, and the Library names it in
spec.trickplay.gpuResourceClaimTemplate or
spec.appearances.gpuResourceClaimTemplate. Each pod of the worker
claims a device from that template. The field is named for the role
of the claim, the GPU the worker decodes on, so a claim for another
kind of device, such as an NPU for the face models, takes a field of
its own.
Both workers need a GPU whose VA-API driver decodes the library’s
video, 10-bit HEVC included, and scales 10-bit frames. A render node
alone does not say what its driver can do. On a fleet of mixed GPUs, a
pod can then land on a GPU that the image’s driver cannot open, such as
an AMD GPU when the image holds the Intel driver. ffmpeg fails on
every video before it reads a frame, and the worker records a miss for
each one, which lasts 30 days. So the template pairs the render node
with a media.liken.sh capability device of the same GPU, and the
scheduler places the pod only on a GPU whose driver states that it
decodes and scales 10-bit video.
Claim a GPU by what it decodes
describes the capability devices and writes the decode-10bit class.
The gpu-render class here selects every render node that liken
publishes:
apiVersion: resource.k8s.io/v1
kind: DeviceClass
metadata:
name: gpu-render
spec:
selectors:
- cel:
expression: |
device.driver == "liken.sh" &&
has(device.attributes["liken.sh"].renderNode)
---
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
name: trickplay-gpu
namespace: media
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu-render
- name: decodes
exactly:
deviceClassName: decode-10bit
constraints:
- requests: [gpu, decodes]
matchAttribute: resource.kubernetes.io/pciBusID
apiVersion: library.liken.sh/v1alpha1
kind: Library
metadata:
name: movies
namespace: media
spec:
trickplay:
enabled: true
gpuResourceClaimTemplate: trickplay-gpu
The appearances worker takes a template of the same shape in
spec.appearances.gpuResourceClaimTemplate. On a fleet where every
GPU’s driver decodes the library’s video, the gpu request alone is
enough.
This template needs a liken release that publishes
resource.kubernetes.io/pciBusID on its render nodes. On an older
release, the claim never allocates, and the worker’s pods stay
Pending.
The operator reads only whether the template exists. When the
Library names a template that its namespace does not hold, the
operator starts no worker of that fact, because the worker’s pods
would stay Pending until the Job’s 24-hour deadline. The
GPUClaimTemplates condition of the Library is then False with
the reason ClaimTemplateNotFound, and its message names the
template:
kubectl -n media get library movies -o jsonpath='{.status.conditions[?(@.type=="GPUClaimTemplates")]}'
The operator watches the templates, so the next worker starts as soon
as you create the template. A template that exists but that no node
can allocate leaves the worker’s pods Pending, and their events say
so.
The Library schema has no render field. An earlier release of the
operator built a template for each worker from that field, named
<library>-trickplay or <library>-appearances, with the Library
as its owner. The operator deletes each such template the first time
it reconciles the Library, and leaves a template you wrote at the
same name.
A worker on several nodes
At the default parallelism, the worker decodes one video at a time. A
large library can decode several videos at once, on several nodes'
GPUs, with spec.trickplay.parallelism or
spec.appearances.parallelism, the number of pods the worker Job
runs at once, from 1 to 16:
spec:
appearances:
enabled: true
parallelism: 3
gpuResourceClaimTemplate: appearances-gpu
Each pod claims its own device from the template that
gpuResourceClaimTemplate names. The pods prefer different nodes.
When fewer nodes offer the device than the worker runs pods, two pods
share a node and its device, because liken publishes a render node
for many claims at once. Every pod goes through the scheduler on its
own, so a node’s taint and the pod’s resource requests apply to each
one. Kubernetes retries a failed pod up to twice on its own video, and
leaves the other indexes alone. The 24-hour deadline is for the whole
Job. The worker counts as running until every index has ended, so
the next worker waits for the last pod.
Two pods can work two episodes of one season at once, and each writes the season folder’s ledger. Each write reads the ledger again just before it replaces the file, and applies the pod’s own entry to what the other pod left, so both attempts stay. The two writes can still collide in the moment between that read and the rename. Each pod asks for a rescan of the series folder, and the operator holds the requests until its next pass, which walks the folder once.
The tile directory is the one Jellyfin reads and writes. So the worker
accepts a directory Jellyfin made first and leaves it alone. The one
exception is a directory older than the file beside it: when a new
file replaced the video at the same path, tiles made before it arrived
place each thumbnail at the earlier file’s times. The worker decodes
the new file and replaces the whole directory. The
scanning guide
says how the walk finds a replaced file. To make the tiles of any
title again, delete its .trickplay directory, and the next walk opens
its gap, and the end of that Job starts a worker.
Appearances
The appearances fact records which credited actor is on screen at
each second of a feature. It runs where spec.appearances.enabled is
set. The fact asks no provider: it matches the faces in the video with
the headshots that the people facts put in .contributors/. A first
pass decodes every frame of every feature, so the fact runs in a
worker Job of its own, named <library>-appearances-<suffix>, in the
same way as the trickplay
worker. The Job that walks
the Library publishes the appearances gap on the bus under
missing/appearances, and each pod of the worker reads one video,
checks it again before it opens it, and asks the operator to rescan
its title folder when it is done.
spec.appearances.parallelism runs the worker on several nodes at
once, as a worker on several nodes
describes.
A feature is in the gap when the probe gave it a length and its title credits at least one actor whose entry holds a headshot. The credits fact credits a series and not each episode, so an episode takes the cast of its series. No source records an episode’s guest stars yet, so a guest star has no headshot to match. A title whose cast has no headshot waits until the headshot fact writes one.
For each video, the worker runs two passes of the appearances tool:
appearances detectdecodes every frame of the video withffmpeg, finds the faces with the YuNet detector in every keyframe and in a frame every second, embeds each face with the SFace model, and writes the detections record.liken/appearances/<file>.jsonlbeside the video. This is the only pass that opens the video. The detector keeps each face it scores at 0.8 or more, and the record marks each keyframe. The record’s first line names the record’s format, the file’s size, and the hash of each model. When a record on the volume names the current format,liken.sh/appearances/detections/v2, the file’s size, and the models of the image, the worker uses it and does not decode again. A record of any other format holds the keyframes alone, so the worker decodes its video again.appearances matchembeds the headshot of each actor that the title folder’s.liken/credits.yamlnames, and names the faces in two passes. The headshot pass names a face for the closest actor when the similarity is at least 0.363, the threshold OpenCV publishes for SFace, and leads the next closest actor by at least 0.05. The film pass then names a face the headshots missed, such as a face in profile or decades younger than the headshot, when at least 2 faces the headshots named in other samples are close to it, all of them name one actor, and they are at least 30% of the faces close to it. The match takes 1 to 5 seconds, and it runs on the CPU.
The worker writes the answer to .liken/appearances.yaml beside the
video, one entry per file:
appearances:
- path: Example Movie (2019).mkv
size: 4831838208
embedder: {name: face_recognition_sface_2021dec, sha256: 0ba9fbfa...}
threshold: 0.363
margin: 0.05
gallery:
- {contributor: .contributors/ad/ada-quill, name: Ada Quill, headshot: found, sha256: 5d2c...}
- {contributor: .contributors/bo/bo-reyes, name: Bo Reyes, headshot: missing}
unmatched:
- {contributor: .contributors/bo/bo-reyes, name: Bo Reyes, headshot: missing}
named: {headshot: 4326, film: 921}
people: 21
attempts:
- path: Example Movie (2019).mkv
at: 2026-10-01T12:00:00Z
result: found
The entry records the inputs of the answer: the size of the file, the
embedding model, the threshold, the margin, and each actor’s headshot
with its hash. named counts the faces each pass named, and people
counts the actors they name. unmatched lists the actors who cannot be
named, because their entry holds no headshot or the detector found no
face in it.
The faces themselves are in .liken/appearances/<file>.matches.json
beside the video, one JSON document of 0.4 to 0.8 MB for a film. Each
face is one observation: the time of its sample in seconds, the face’s
place in that sample’s line of the detections record, the actor, the
pass that named the face, the similarity to the actor’s headshot, and
the closest other actor’s similarity. A face named by film can be
under the threshold. The ledger holds only the counts, because a walk
reads every ledger of a folder on each pass, and the faces of one film
would make the ledger entry about 700 KB. A failed pass is an error
attempt whose reason holds the tool’s own error text, and the worker
tries the file again after a day.
The match also writes .liken/appearances/<file>.spans.json beside
the video, for the screens. It holds spans, not samples: each run of
samples that names the same actors, with the actors left to right in
the order they stand in the picture, and each actor’s name, character,
headshot, and dates. Each sample stands for the time halfway to its
neighbours, and never past the next keyframe, where a cut can fall. The
actors named last hold through samples that name nobody for up to 10
seconds, and two samples in a row with no face end the span. When a person plays a video whose answer is
found, the screen’s browser names the spans file and the library’s
.contributors/ in the Play. The display then shows who is on screen
when the film pauses. The
tool’s README
gives the file’s shape.
A found answer stands until a new file replaces the video at its path.
The answer does not change when a headshot or the credits change. To
match a title again, delete its .liken/appearances.yaml. The next
walk opens the gap, and the worker uses the detections records on the
volume and runs only the match. To match every title again, set
spec.refresh.appearances to the current time, as the
When a fact asks again
section describes
for every fact. A refresh also decodes again each video whose record
is of an earlier format, and the match then rewrites its ledger entry
and its spans file in place.
kubectl -n media get jobs -l library.liken.sh/library=movies,library.liken.sh/worker=appearances
The worker runs on the library-operator-appearances image, which
carries the tool, Intel’s OpenVINO runtime, Intel’s OpenCL runtime for
the GPU, and the two models from
OpenCV Zoo
: YuNet under the MIT
license and SFace under the Apache 2.0 license. The operator names the
image at its own tag, and APPEARANCES_IMAGE on the operator’s
Deployment names another. The container requests one core of CPU and
may take 1536Mi of memory. The largest measured run held 740 MB in
the tool and ffmpeg together, for a 4K file decoded in software. On a
laptop, the detect pass took 1.5 to 4 minutes for a 1080p film and 6
minutes for a 4K film through VA-API, and 34 minutes for the 4K film in
software. The worker allows each detect pass 6 hours.
With no template, the worker decodes and runs the models on the CPU.
spec.appearances.gpuResourceClaimTemplate names a
ResourceClaimTemplate, as A worker on a GPU
describes for both workers. With the claim, ffmpeg decodes and scales the video on
the render node through VA-API, and OpenVINO runs the models on the
Intel GPU. Some GPUs decode a format that their video processor cannot
scale, such as 10-bit HEVC on an older Intel GPU. The worker tries the
GPU’s scale on the first 2 seconds of each video, and where that fails,
the GPU decodes and the CPU scales. A file the render node refuses to
decode is decoded again in software. OpenVINO compiles the models for the GPU when the worker
starts, which took seconds in the measurements, and it keeps the
compiled kernels in an emptyDir that the detect and match passes of
the pod’s one video share.
Checking the faces by hand
appearances review shows a person what the match named, on a copy of
a title folder on a workstation. It runs no model and writes nothing
into the folder. Never run the tool on the library itself. The copy
needs the title folder with its .liken directory, and the
.contributors/ entries its credits name, at the same paths under a
common root. The
tool’s README
says how to build it and how to install the runtime.
appearances match "Movies/Example Movie (2019)"
appearances review "Movies/Example Movie (2019)" --sheets 24 --play
--play opens mpv with a chapter for each span and a box around each
face: green for a named face, yellow for a face near the threshold or
the margin, and red for a face no actor is close to. --sheets 24
writes contact sheets of the faces named for each actor, weakest
first. --threshold and --margin show the result of other values
with no pass over the video. For an episode, run the match on the
season folder and name the series’ credits:
appearances match "Series/Example Show/Season 01" --credits "Series/Example Show/.liken/credits.yaml"
Identification
The identity fact asks TMDb for the folder’s title and runs a fixed
sequence of tests: the title, then the year, then a year on either side,
then, for a series, the episode names, then the runtime within five
minutes. The episode test reads the episode titles from the file names
of the folder’s first season. It keeps a candidate whose season on TMDb
has at least two of them and at least half. So a series folder named
with the title alone identifies without an .nfo file. One survivor is the
answer, and its reason is recorded. Several survivors become candidates
in .liken/identity.yaml, and the title counts in status.waiting
until a person names the right uniqueid in the .nfo. A title no
provider can name counts in status.unresolved.
After identifying a title through TMDb, the identity phase requests its IMDb
and TVDB ids and writes any returned ids into the .nfo file. OMDb uses
the IMDb id to look up the title. Fanart.tv uses the TMDb id for a
movie and the TVDB id for a series. These ids identify the title;
OMDb and Fanart.tv still require their own API keys, configured through
each provider’s secretRef.
The write rule
Each fact owns a fixed group of elements in the .nfo file and writes
nothing outside it. The overview fact owns the plot, the tagline,
the genres, the studios, the premiere date, and the runtime. Each
rating owns its one element. credits owns the actors, directors,
and writers. Before a fact writes again, it hashes the group as it is
now and compares that hash with the one it recorded after its last
write. If the hashes differ, another writer changed that group. The
fact records a fight and does not write the group. status.fights
counts the fights.
An art file that already exists is never replaced. The fact records it as answered and downloads nothing. The one exception is an episode’s thumbnail that is older than the episode’s file, where a new file replaced the video at the same path. That thumbnail is a frame of the earlier file, so the fact downloads the still again and writes it over the old one.
Every file lands whole: the writer fills a temporary beside it and
renames the temporary into place. Two clusters can mount one library,
each with its own Library over the same root, and their writers can
reach one file at once. A writer that changes a file it read, an
.nfo file, a .liken ledger, or a person’s contributor.yaml,
reads the file again just before its rename. Where another writer
changed the file in between, it applies its change to what that
writer left, so neither change is lost. The two clusters can still do
the same work twice. The check keeps each other’s writes, and it does
not divide the work between them.
People
The credits fact writes each title’s cast and crew into
.liken/credits.yaml, and it gives each credited person one entry in
.contributors/ at the library root. A credit looks for its entry by
id first. When the catalog holds an entry with one of the credit’s
ids, the credit names that entry, whatever the spelling of the name.
A credit that holds an IMDb id and no TMDb id asks TMDb for the TMDb
id first, when the Library’s sources name a Ready tmdb provider.
When no id finds an entry, the credit uses the entry at the slug of
the name. A second person of the same name gets the slug with an id,
such as nora-vance-tmdb-992.
A credit from IMDb’s datasets names a person by the IMDb id alone. So
it finds the entry a TMDb credit wrote for the same person by that id,
or through the TMDb id that TMDb’s find call gives for it, and the
person keeps one entry. An entry the datasets created holds the IMDb id,
and contributor.ids fills its TMDb id, its biography, and its
headshot the same way.
The contributors container fills each entry. contributor.ids
writes the birth date, the death date, and the person’s ids in other
databases into contributor.yaml. For an entry that holds only an
IMDb id, it asks TMDb for the TMDb id first. contributor.biography
and contributor.headshot write biography.txt and headshot.jpg
beside the entry where no file of that name exists.
Two entries can hold one id, for example when two providers spell one
name two ways. After it writes the ids, contributor.ids merges each
group of entries that share an id into one entry:
- The entry at the slug of its own name, with no id suffix, stays. If no entry of the group is at such a slug, the entry with the most ids stays.
- The entry that stays gets every id of the group.
biography.txtandheadshot.jpgmove to it where it has no file of that name. - Each other entry keeps a
contributor.yamlwith one field,mergedInto, the path of the entry that stays. The catalog shows no person for such an entry. - The next
Jobof theLibrarymoves each credit that names a removed entry to the entry that stays. The nfo phase moves the credits, and it can end before the merge, so the move waits for thatJob. The contributors phase of the sameJobthen deletes each removed entry that no credit names.
The merge leaves a group whole in two cases, and
.liken/contributor.merge.yaml in each entry of the group records
the reason. A held attempt means a person edited a
contributor.yaml of the group, and status.fights counts each entry
of it. A conflict attempt means two entries hold two different ids
in one scheme, so they are two people and one of them holds a wrong
id. The merge asks about the group again after thirty days. To hand
an edited entry back to the merge sooner, delete
.liken/contributor.ids.yaml and .liken/contributor.merge.yaml in
that entry. The next walk opens the gap, and contributor.ids then
reads the file as its own.
Trailers
The trailer fact records links, not files. It asks every source
that serves it and writes what it finds to .liken/trailer.yaml
beside the title, one entry per video: the provider, the provider’s
own key, the site the video plays from, the page to watch it on, its
name, its kind (trailer, teaser, or spot), its language, its
resolution where the provider states one, and a score from 1 to 100
with a one-line reason. The catalog’s trailers table holds the same
rows.
The score ranks the videos of one title, so a screen can take the first.
A video keyed by the title’s own id, as TMDb’s are, starts at 100. A
video found by a search starts at 90 when its name has the same title
and year. It starts at 60 when the name has the title and no year at
all. A search result whose name has another title or another year is not
recorded. Then the kind takes points off, a teaser less than a TV spot.
A language the household does not prefer takes points off, and so does a
TMDb video that is not marked official. The reason says which of these
applied. Clips, featurettes, and other extras are not recorded at all.
Within one provider, videos with the same name collapse to the best one,
and at most five videos per provider are kept for a title. An Internet
Archive item is kept only when its own video runs eight minutes or less,
because the movie_trailers collection holds whole films beside the
trailers.
The preferred languages are the library’s spec.languages, then the
household’s audioLanguages from the media operator’s
MediaPreferences, and en when neither names any.
A trailer file beside the title still plays as before.
Intro and credits marks
The marks fact records where a video file’s intro, recap, credits,
and preview are. It asks every source that serves it about each main
video of an identified movie, and of each episode of an identified
series, once the probe fact has measured the file’s length. It writes
what it finds to .liken/marks.yaml in the folder that holds the file,
the folder whose .liken/probe.yaml records the same file:
marks:
- path: A Series - S01E02.mkv
kind: intro
end: 107000
source: theintrodb
- path: A Series - S01E02.mkv
kind: intro
start: 7007
end: 106482
source: theintrodb
- path: A Series - S01E02.mkv
kind: credits
start: 3253000
end: 3316000
source: theintrodb
Each entry is one span: the file, the kind, the start and the end in
milliseconds from the start of the file, and the provider that answered.
An absent start is the start of the file, and an absent end is the
end of the file. A provider can answer several candidates for one kind,
from several submissions or several releases of the work, and the fact
records every one exactly as the provider answered it. It chooses none.
The catalog’s marks table holds one row per span, and the media
browser sends every span to the Play, where the player in
media-operator reads them and offers the skip.
TheIntroDB reads the file’s length from the ask, and it answers the spans of the release whose length is closest, such as the theatrical cut or the extended one. IntroDB reads no length. A file that holds two episodes is asked about neither, because each provider places an episode’s spans in that episode’s own file.
The community adds marks for a new episode or film in the days after it comes out, so the fact asks again about a new work sooner than about an old one. The release date the catalog holds for the movie or the episode sets the wait after a find or a miss:
| Released | Asked again after |
|---|---|
| in the last 7 days | 1 day |
| 8 to 90 days ago | 7 days |
| more than 90 days ago, or no date | 30 days |
A file that holds two episodes takes the later of their dates. An error waits one day, whatever the date.
The table applies only where every provider the Library names
answered. Where one provider failed, or was not asked, the attempt
records the result partial. The file keeps the spans the other
providers answered, keeps the spans the silent provider answered
before, and waits one day, so that provider is asked again the next
day.
A provider that answers 429 waits for the reset its headers name and
asks again. When a provider has spent its allowance for the day, the
reset is hours away, and the fact asks that provider nothing more in
this run. Every later file of the run is partial, and the run stops
when every provider has spent its allowance. A file that every provider
answered keeps the window in the table, so the next day’s allowance
goes to the files the spent provider missed and not again to the ones
it answered.
Neither Jellyfin nor Kodi reads these marks. Jellyfin keeps its media
segments in its own database, and Kodi’s .edl file is a different
format, so the fact writes no file other than its ledger.
Trailer files
The trailerfile fact pulls one video file per title. It is off unless
spec.trailers.enabled is true, and it pulls nothing until the
trailer fact has recorded links.
The file lands at <title>/trailers/<name>.mp4. The name is the
trailer’s own name with every character a file name cannot hold taken
out, and its length is capped.
The fact takes the highest-scored trailer whose site the operator can fetch from. Today those sites are the Internet Archive and PeerTube, never YouTube. From that trailer’s files it takes the tallest file that is no taller than the title’s own feature, or the shortest above it where none fits.
The fact never pulls for a title that already holds a trailer file anywhere under its folder. A trailer a person placed by hand stays, and the fact records nothing.
Every pull is remuxed to MP4 and checked with ffprobe before it
lands. The check requires a video stream and a length between 10
seconds and 8 minutes, so a TV spot passes and a whole film does not.
A file that fails the check never reaches a name the walk reads, and
the attempt records the error.
A trailer is tens of megabytes per title, so a library of any size adds gigabytes to the volume the first time this fact runs.
The trailer-files phase of the Library’s Job runs the fact, and
two titles pull at once inside it. The phase starts no new title
fifteen minutes into a run, and the next Job goes on with the rest. The fact takes no spec.refresh.
IMDb ratings and credits from the datasets
IMDb publishes its datasets as gzipped files at datasets.imdbws.com
for personal and non-commercial use. No liken image carries them, so
each cluster downloads them from IMDb. IMDb’s terms require this credit
where the data is shown:
Information courtesy of IMDb (https://www.imdb.com ). Used with permission.
The nfo container reads its whole rating.imdb gap first. Then it
reads title.ratings once, from start to end, and keeps only the rows
of the titles in the gap, so one read serves one title or ten thousand.
The read starts when the container starts, and the other facts run
while it works. A run with no rating.imdb gap sends no request.
A movie or a series is found by the IMDb id in its .nfo file. An
episode is found by its own IMDb id where its .nfo file names one.
Otherwise the container reads title.episode once and finds the
episode by its season and episode numbers under the series’ IMDb id. It
records the id it found in .liken/rating.imdb.yaml, so a later run
does not read title.episode again for that episode. A file that holds
two episodes gets no rating, because its ledger entry names one episode.
The credits read two files in sequence. The container reads
title.principals once and keeps the rows of the titles in its
credits gap, then reads name.basics once for the names of the
people those rows name. It does not read name.basics when every one
of those people already has an entry in .contributors/, because the
entry holds the name. Decompressing and reading the two files took
about 32 seconds on a cloud host in September 2026, however many titles
they fill, and the downloads add to that on the first run of a node. actor, actress, and self
rows are the cast, in IMDb’s order, with the characters as the role.
director and writer rows are the crew. The other categories, such
as producer and composer, have no part in the credits and are not
written, and neither are archive_footage and archive_sound. The
credits are not asked for again on a timer. Set credits in
spec.refresh to ask again.
The rating is written into the .nfo file as OMDb writes it: the
imdb rating, out of 10, with the vote count. The container writes the
file only when the rating at one decimal changed. A change in the vote
count alone writes nothing, so a library’s .nfo files do not all
change every month. The attempt records the Last-Modified time of the
title.ratings copy it read.
A rating from the datasets is asked for again after 30 days, but only
when IMDb has published a newer title.ratings than the one the
attempt read. A rating OMDb wrote counts as older, so a library that
moves from omdb to imdb moves each title within 30 days. A rating
that another tool, such as Radarr or Jellyfin, wrote into the .nfo
file has no attempt. The next nfo container that reads the datasets
reads that rating once and records an attempt, and it writes the file
only when the rating at one decimal differs. The gap count in the
status does not include these ratings, so they start no Job of their
own. While the
provider’s Stale condition is True, IMDb has published no newer
file, and the operator starts no Job that fills gaps for these ratings
alone.
With a per-node StorageClass in the cluster, the operator keeps the
files on the claim <provider>-datasets, which the nfo container
mounts. Each node’s copy fills the first time a run on that node needs
a file. Every later run asks IMDb with the copy’s ETag, gets 304 Not Modified while the file is unchanged, and reads the copy from disk. A
new version replaces the copy only after the whole file arrived and its
gzip checksum is correct. Two runs on one node download a file once:
the second waits for the first and reads its copy. When the claim is
full or cannot be written, the container reads the file from IMDb and
logs the filesystem’s error. With no per-node class, every run reads
the files from IMDb.
kubectl -n media get metadataprovider imdb -o jsonpath='{.status.imdb.datasets}'
kubectl -n media logs job/<job> -c nfo | grep -E 'title.ratings|title.principals'
When a fact asks again
A miss lasts for thirty days and an error for one day, then the fact
asks again. The marks fact asks about a new work sooner, as the
section above describes. An attempt made before a title’s release date
lasts only until that date. An attempt also stops counting when the walk finds that a new file
replaced the video at its path, or that the file or directory the
attempt wrote is gone. The
scanning guide
describes both. To ask one fact again for every title now, set its
time in spec.refresh:
spec:
refresh:
overview: "2026-09-06T00:00:00Z"
Every attempt of that fact before the time no longer counts. A refresh
starts a Job that fills gaps, without a walk, as soon as the
Library has no other Job running, and a refresh set while a Job
runs starts another one after it. The fact
rewrites its own files and rows in place, and nothing is deleted.
The appearances key works the same way: the reopened videos are in
the gap the worker reads after the Job that fills gaps, and the
worker matches them again. The worker takes the videos whose detections
record is on the volume first, because each of them needs only the
match.
kubectl liken library reenrich movies writes that field for you and
asks every fact again. Add --only overview to reopen one fact, and
-n to select the namespace. The command reads your kubeconfig, and
in bash it completes the library names.
The same map takes one key that is not a fact: scan asks for a full
walk of the library, and kubectl liken library rescan movies writes
it. The scanning guide
describes what it does.
4. The Jellyfin handover
Jellyfin reads the .nfo files and the art that this operator writes, under the
same names its own scraper uses. Turn off Jellyfin’s
SaveLocalMetadata for a library this operator enriches. With it
on, two writers change the same .nfo files, and the fight check leaves
the file to Jellyfin.
Reading progress
kubectl -n media get library movies -o jsonpath='{.status.gaps}'
kubectl -n media get library movies -o jsonpath='{.status.waiting} {.status.unresolved} {.status.fights}'
kubectl -n media get library movies -o jsonpath='{.status.conditions[?(@.type=="Sources")]}'
status.gaps counts, per fact, the rows still to fill. Between walks,
the operator starts a Job that fills gaps with the phases whose counts
are above zero, once a refresh time or a source provider that turned
Ready has given them work since the last Job started. trailer-files
runs in such a Job while its count is above zero, because each run
stops at its time limit. The trickplay and appearances counts
start no Job that fills gaps: each starts its fact’s worker when a
Job of the Library ends.