
# Scanning

A scan walks a library's root and writes what it finds into the
namespace's catalog. This guide describes what the walk reads, when
it runs, how it removes what is gone, and what happens when a
`Library` is deleted.

## What a scan reads

The scanner reads the layout Kodi and Jellyfin read, so a volume those
players already organize needs no change.

A `Library`'s root holds title folders and grouping folders. It never
holds a title of its own. The walk reads titles only in the folders
below the root, so it catalogs no video file that sits directly in the
root, and none in a `trailers` or other extras folder at the root. Put
each film in a folder of its own.

### Movies

One folder per title. A folder is a title folder when it holds
`movie.nfo` or a video file. Any other folder is a grouping folder,
and the walk descends through it, up to eight levels deep:

    movies-pvc/
      Action/
        Example Movie (2019)/
          Example Movie (2019).mkv
          movie.nfo
          folder.jpg
          fanart.jpg
          Example Movie (2019).trickplay/
          Extras/
            Making Of.mkv
          Trailers/
            Example Movie (2019)-trailer.mkv

A readable `movie.nfo` with a title names the movie. Without one, the
folder name is parsed as `Title (Year)` or `Title [Year]`, or cut at
the first release token such as `bluray` or `x264`. A folder with no
`.nfo` file and no year is counted in `status.unidentified` and cataloged
under its folder name.

### Series

One folder per series, directly under the root. Episodes are the files
in the series folder and in its season folders, one level down:

    series-pvc/
      Example Series/
        tvshow.nfo
        folder.jpg
        Season 02/
          season02-poster.jpg
          Example Series - S02E05.mkv
          Example Series - S02E05.nfo
          Example Series - S02E05-thumb.jpg
          Example Series - S02E05.en.srt
        Specials/
          Example Series - S00E01.mkv

`Season NN` is season N, and `Specials` is season 0. An episode's
number comes from a marker in its file name, `s02e05` or `2x05`, and a
range such as `s04e10-e11` names two episodes in one file. Both play
the file from the start, because nothing on the volume marks where the
second begins. The season comes from the folder first, then from the
marker.

### Extras and file kinds

A folder named `extras`, `featurettes`, `trailers`, `behind the
scenes`, `deleted scenes`, `interviews`, `scenes`, `shorts`, `clips`,
or `other` beside a feature or a season is read one level deep. Videos
under `trailers` are trailers, and the rest are extras.

A folder with one of those names that holds video files is an extras
folder wherever it is, at the library root or inside a grouping folder.
The scanner reads no title from it, and nothing under it is cataloged.
A folder with one of those names that holds only folders is a grouping
folder. So a genre folder named `Shorts` is read, and the titles under
it are cataloged. A title whose own name is one of those words has its
year in its folder name, `Trailers (2016)`, which is not the bare
word.

Every file gets a row. The scanner classifies each one as `video`,
`audio`, `subtitle`, `image`, `metadata`, `trickplay`, or `other`
from its name, its folder, and one `stat`. It opens no file to
classify it. A subtitle's language is the tag in its name, and `hi`
after a language tag marks it hearing-impaired. Dot-named entries,
`Thumbs.db`, `desktop.ini`, and the trash and service directories of
common NAS systems are skipped.

### .nfo files and art

The scanner reads `movie.nfo`, `tvshow.nfo`, and the `.nfo` beside
each episode, leniently, because Jellyfin writes bare ampersands in
URLs. It reads the title, the year, the plot, the genres, the people,
the ratings, and the provider ids in `uniqueid` elements, with
`imdbid`, `tmdbid`, and `tvdbid` as fallbacks. The scanner writes the
genres to the catalog in the spelling the enricher uses, so a `.nfo`
file that holds `Sci-Fi` shows `Science Fiction` in the browser. The
file itself does not change, whether another program or an earlier
enricher wrote it.

Art uses Kodi's names: `poster.jpg`, `fanart.jpg`, `clearlogo.png`,
`clearart.png`, `banner.jpg`, `landscape.jpg`, and `disc.png` in the
title folder, `season02-poster.jpg` beside `tvshow.nfo`, and
`<episode>-thumb.jpg` beside the episode. The scanner also accepts
`folder.jpg` and name-prefixed forms such as `<title>-poster.jpg`.

### The `.liken/` directory

In each title folder, a dot-named directory holds the data that the `.nfo` file has
no element for: one YAML file per fact, named for the fact. `identity.yaml`
holds the provider ids, or the candidates left for a person to choose
from. `arrival.yaml` holds when each video file was first seen.
`probe.yaml` holds what the probe read of each file, its size included.
Every other `<fact>.yaml` holds what that fact wrote, which provider
answered, and its attempts. One file per writer lets the phases of a `Job` run at
once on a network mount with no locks. The scan reads these files and
never writes them.

Every writer writes its file under a partial name beside the final name
and renames it into place, so a reader never reads half a file. The
operator's partial names carry `.liken-tmp-`, and the appearances tool's
carry `.partial-<host>-<pid>`, in `.liken/` and in `.liken/appearances/`.
A writer that is stopped before the rename leaves its partial file
behind. The walk lists each `.liken/` directory, and the
`.liken/appearances/` directory where one exists, and names each partial
file that no writer has changed for 24 hours. The scan mounts the volume
read-only, so it hands those names to the `Job`'s close container, which
removes each file and logs one line for it. The longest writer, the
appearances tool's detect pass, stops after 6 hours and changes its file as
it writes, so the walk never names a file that a live writer holds.

The library root holds no title, so no fact writes a `.liken/`
directory there. The close container of every `Library` `Job` removes
a `.liken/` directory at the root, with every file in it.

### `.contributors/`

At the library root, one directory per credited person, sharded by
the first two characters of the person's slug. Each holds
`contributor.yaml` with the name and the provider ids, and, once the
contributors phase fills them, `biography.txt` and `headshot.jpg`. The walk
reads this directory after the titles. It is the one dot-named
directory the walk enters. An entry whose `contributor.yaml` holds
only `mergedInto` is one that a merge of two entries removed. The walk
records the merge and no person for it. The
[enrichment guide](/docs/guides/enrichment/#people) describes the
merge.

### A file replaced at the same path

A download manager that upgrades a title imports the new file under
the old name, and it often gives the new file the old modified time.
So the walk compares each video's size with the size in its
`probe.yaml` record. A file of another size is a new file:

- The walk stops counting the `probe`, `trickplay`, `episode-thumb`,
  `marks`, and `appearances` attempts at that path, and it drops the file's marks from the
  catalog. The file's length and codecs are empty until the probe reads
  the new file.
- The tile directory and the thumbnail beside the file were made from
  the earlier file, so the walk no longer counts them as this file's.
  The trickplay worker and the `art` phase make them again and replace
  the old ones. A tile directory or a thumbnail made after the new file
  arrived, by Jellyfin or by any other writer, is this file's, and both
  keep it.
- `arrival.yaml` keeps the path's first arrival, so an upgrade does not
  put an old title back on the recently added rail.

The `trickplay`, `marks`, and `appearances` gaps open once the probe
has measured the new file, because the three facts need its length. A new modified time on a
file of the same size opens only the probe.

A found attempt also needs its output. When the walk finds no tile
directory beside a video whose `trickplay` attempt found one, or no
thumbnail beside an episode, poster, season art, trailer file, headshot,
or biography that its fact recorded, that fact's gap opens. So a person
who deletes an output gets it back on the next walk or webhook rescan.

The walk logs each replaced file with both sizes, and one count of what
it opened:

    replaced file path:6f1c2a90d3b4: 3345988372 bytes in the probe record, 3362396196 on the volume
    reopened the facts of 16 replaced files and 32 missing outputs

## When a scan runs

Every scan runs in a `Job` the operator creates for the `Library`, and
only one `Job` of a `Library` runs at a time. The operator starts a
walk `Job` when one of these asks for it and no other `Job` of the
`Library` is unfinished:

* `spec.scan.schedule`, a cron expression in the form a `CronJob`
  takes, once an hour by default. It is in UTC unless it starts with a
  `CRON_TZ=` prefix. A walk is due when a time in the
  schedule has passed since the last full walk started. A `Library`
  that has never been walked is due at once.
* A request in `spec.refresh.scan`. `kubectl liken library rescan
  movies` writes the current time there. A walk that starts at or
  after the time answers the request, so asking again is a matter of
  setting a later time. It is the same map the
  [enrichment guide](/docs/guides/enrichment/) uses for facts.
* A [webhook](/docs/guides/webhooks/). The walk reads only the folders
  the webhooks named, and a webhook that named no folder asks for a
  full walk.

A walk that is due while another `Job` of the `Library` runs waits for
that `Job` to finish, and the folders webhooks name in the meantime
wait with it. The next `Job` walks all of them.

    kubectl -n media get jobs -l library.liken.sh/library=movies,library.liken.sh/worker=walk

A walk `Job` is named `<library>-walk-<suffix>`. It runs the `scan`
container beside every phase the `Library`'s sources serve: the probe,
the arrival fact, identity, the `.nfo` facts, the art, the trailers,
the marks, the people, and the trailer files. All of them start
together, and each phase works on a title as soon as the walk and the
phases before it have written that title's rows. Trickplay and the
appearances each run in a worker `Job` of their own after the walk
`Job` ends. The
[enrichment guide](/docs/guides/enrichment/#3-what-the-phases-do)
describes the phases. The `scan` container runs a person's own image
when the kind's settings block names one, and it always mounts the
library volume read-only.

The walk writes a `runs` row under the `scan` worker when it starts,
and again when it finishes. A walk of folders writes its row under the
`rescan` worker, so the full walk's counts stay beside it. The `Job`'s
`close` container writes the `enrich` row after the last phase and
waits until a catalog pod confirms it. So a `Job` that completed is a
`Job` whose rows reached a durable copy of the catalog.

Every `Job` of a `Library` runs its catalog agent on the `Library`'s one
catalog claim, `<library>-catalog`. The operator's rule of one `Job` at
a time is what keeps two agents off one database. The trickplay worker
runs no agent and mounts no catalog claim, so it runs beside these
`Job`s and the rule leaves it out. On a per-node class
the claim is also `ReadWriteOncePod`, so the scheduler keeps a second
pod of the claim `Pending` while the first one runs.

## A Job that does not start or that fails

Only one `Job` of a `Library` runs at a time, so a `Job` whose pod
cannot start holds back every walk and every phase of that `Library`.
The `Library`'s status names such a `Job`. When the pod of a `Job`
stays `Pending` for five minutes, the phase is `Blocked`, and the
`Ready` condition is `False` with the reason `JobNotStarted`:

    $ kubectl -n media get libraries
    NAME         KIND         TITLES   ITEMS   FILES   WAITING   SOURCES   STATUS    READY   AGE
    franchises   franchises   12       12      0       0                   Blocked   False   19d

    $ kubectl -n media get library franchises -o jsonpath='{.status.conditions[?(@.type=="Ready")].message}{"\n"}'
    the pod franchises-walk-dlov5hvq2ryn-gklk6 of the Job franchises-walk-dlov5hvq2ryn has not started: GitVolumeRefused: readOnly: a claim on this driver has to be mounted read-only; set readOnly: true on the pod's persistentVolumeClaim volume

The text after the `Job`'s name is from Kubernetes. When no node can
take the pod, it is the scheduler's reason, `Unschedulable`, and its
sentence. When a node took the pod, it is the newest `Warning` event
about the pod, for example a volume that a CSI driver refused or an
image that the kubelet cannot pull. `kubectl describe pod` shows every
event of the pod.

Repair what the message names. A `Job` keeps the pod spec it was
created with, so when the repair is in the `Library` or in the
operator, delete the `Job`. The next pass then creates the `Job` that
is due from the `Library` as it is now:

    kubectl -n media delete job franchises-walk-dlov5hvq2ryn

A trickplay worker holds back no other `Job`, so the `Ready` condition
names neither a worker whose pod has not started nor a worker that
failed. Read its pods by its worker label:

    kubectl -n media get pods -l library.liken.sh/library=movies,library.liken.sh/worker=trickplay

Every `Job` of a `Library` that runs a catalog agent has a deadline of
two hours in `activeDeadlineSeconds`, and the time its pod stays
`Pending` counts. The longest healthy `Job` is shorter: the trailer
files start no title after 15 minutes, and one trailer file takes at
most 20 minutes. At the deadline, Kubernetes fails the `Job` with the
reason `DeadlineExceeded`, so a `Job` whose pod never starts holds back
the next `Job` of the `Library` for at most two hours.

A `Job` also fails when three of its pods fail, with the reason
`BackoffLimitExceeded`. After a `Job` fails, the phase is `Failed`,
and `Ready` is `False` with the reason `JobFailed`, until a later `Job`
of the `Library` succeeds. The message names the `Job` and the reason.
The operator starts the next `Job` at once after the first failure.
After each failure that follows, it waits 10 seconds, and it doubles
the wait up to 5 minutes. A `Job` that succeeds resets the wait. A
failed `Job` stays for an hour, so its logs can be read:

    kubectl -n media logs job/franchises-walk-dlov5hvq2ryn --all-containers

## Mark and sweep

Each full walk has an epoch. The walk marks every id, path, and link
it reads with that epoch. A prune pass then deletes every row of this
library the epoch did not mark, in batches of five hundred.

Two guards keep a bad walk from emptying a library. A walk that could
not read every directory, or that found less than half of what the
catalog holds, is incomplete. It writes what it read and prunes
nothing, and the log reports it:

    incomplete walk: could not read the whole volume, keeping the last counts

A prune whose epoch marked nothing at all is refused as an error.
`status.removedLastSweep` reports what the last sweep removed, so a
mass delete is visible without a shell.

The walk itself runs eight workers over a shared pool of directories.
That keeps a network volume busy without a burst large enough to slow
a player.

## Deleting a Library

A `Library` has a finalizer, and deleting it starts a departure that
removes its rows from the namespace's catalog. The operator starts no
new `Job` for a deleting `Library`, and it waits for any `Job` of the
`Library` to finish. Then it runs a cleanup `Job` named
`<library>-cleanup` on the `Library`'s catalog claim, which deletes the
rows in batches through its own catalog agent. The `Job` then writes
its own `runs` row and waits for a catalog pod to confirm it, the way
every worker `Job` does. After the confirmation, it deletes that row
too, and waits for a catalog pod to hold the delete, so the catalog
keeps no row of the departed `Library`. The finalizer clears when the
cleanup `Job` succeeds.

While this runs, the phase is `Departing`, and the `Departing`
condition names the step: `ScanRunning` while a walk `Job` runs,
`EnrichRunning` while another `Job` of the `Library` runs, `Sweeping`
while the cleanup `Job` runs, or `Blocked` when the cleanup `Job`
keeps failing or the namespace holds two `Catalogs`. There is no timeout. The operator
reports the blocker for as long as the object is deleting.

A namespace with no `Catalog` releases at once, because nothing there
holds the rows. A library whose own catalog claim is already gone gets
a fresh, empty one for the cleanup `Job`. That `Job`'s agent receives
the rows over gossip and then sweeps them.

## Reading progress

    kubectl -n media get library movies -o jsonpath='{.status.phase} titles={.status.titles} unidentified={.status.unidentified} waiting={.status.waiting} gaps={.status.gaps}{"\n"}'
    kubectl -n media logs -l library.liken.sh/library=movies,library.liken.sh/worker=walk -c scan --tail=100

A finished walk logs its counts:

    walk complete: 128 titles from 131 folders, 3 unidentified, 0 removed, in 41s

`status.unidentified` counts the folders cataloged by name.
`status.waiting` counts the titles a provider returned candidates for,
and the identity phase does not retry those until a person names the
right `uniqueid` in the `.nfo`. `status.gaps` counts, per fact, the
rows the phases still have to fill, including the facts of a replaced
file and the outputs that are gone.

