The catalog
Scanners write to the catalog, and every screen reads from it. The
catalog is a SQLite database that Corrosion
replicates. Each namespace has one catalog cluster. Its catalog pods
are the long-running members, and every pod that reads or writes the catalog
runs another member. This guide describes the Catalog resource that
creates the cluster, how a screen gets its copy, and how the counts
reach a Library’s status.
The Catalog resource
One Catalog per namespace creates the catalog pods. It sizes every
catalog claim in the namespace and owns the Service that the
cluster’s members use to find one another:
apiVersion: library.liken.sh/v1alpha1
kind: Catalog
metadata:
name: media
namespace: media
spec:
storage:
size: 1Gi
storageClassName: local-path
screens:
storageClassName: local-path
artCache:
size: 2Gi
Every member holds the whole namespace’s catalog, because the cluster
gossips every row to every peer. spec.storage.size sets the size of
new catalog claims for the durable replicas, each Library’s Jobs,
and screens. spec.storage.claimName names an existing claim for the
durable replicas in place of the one the operator provisions.
Catalog
describes every field.
The listing shows the storage size, requested replica count, and readiness:
$ kubectl -n media get catalogs
NAME SIZE COPIES READY AGE
media 1Gi 1 True 3d
Ready follows the catalog replica pods and the storage configuration.
A replica that cannot run keeps the catalog from becoming ready;
a screen that is down does not. status.replicas.catalog and
status.replicas.progress report the ready and requested counts for
each store. status.members lists the catalog cluster’s member pods,
and status.screens lists each screen with its claims, node, and phase.
The catalog pod
spec.storage.replicas sets the number of durable catalog replicas
and defaults to one. Their pods are named <catalog>-catalog-0,
<catalog>-catalog-1, and so on, where <catalog> is the resource’s
name. Even a single replica has the -0 suffix.
Every replica runs a catalog container with the Corrosion agent and
a confirmer container. The agent is a native sidecar in
spec.initContainers. The confirmer checks that worker writes reached
this replica, as described below. Replica zero also runs a reporter
container, so it has three running containers; other replicas have two.
The reporter reads its local agent and publishes one retained report
per Library on the bus. It rebuilds the report whenever the catalog
changes, at most once a second. These reports supply a Library’s
counts, gaps, and runs. The reporter counts each gap with the
Library’s spec.refresh times, which the operator publishes on the bus,
so an edit of spec.refresh changes the gaps with no restart of the
catalog pod. The reporter has no Kubernetes credential.
The operator alone writes status.
All durable replicas mount the same claim, named <catalog>-catalog
unless spec.storage.claimName names an existing one. With a per-node
storage class, each node has a separate directory for that claim.
The operator provisions its claims as ReadWriteMany on that class
and places replicas on different nodes. It provisions ReadWriteOnce
claims on other classes. If more than one replica is requested on a class
whose provisioner is not per-node.liken.sh, the operator runs one
replica and reports Ready: False with reason ClassNotPerNode.
A durable replica on a node that has not been Ready for more than ten
minutes cannot move, because its claim is bound to that node. The
operator deletes the pod, and its claim on a class that binds a claim
to a node, so the next pass creates the replica on another node. It
posts a StoreCopyHealed Warning Event on the Catalog that names
the replica and the node. Each change of the Ready condition posts an
Event too, with the condition’s reason and message. ManyCatalogs,
ClassNotPerNode, and PodFailed are Warnings.
The Catalog also creates a separate Corrosion cluster for playback
progress. spec.progress.replicas controls its replica count. Progress
has separate storage because a catalog rescan cannot recover it.
The agent’s API binds to loopback, so nothing on the pod network can
reach it. The pod’s probes run SELECT 1 through the agent’s own
binary inside the container. The startup probe allows ninety seconds
for the agent to open its database.
How the members find each other
The operator writes a headless Service named catalog in the
namespace, on UDP port 8787, and writes its EndpointSlice itself.
The slice holds every pod in the namespace that has the member
label: the catalog replica pods, the running Job of each Library,
a running cleanup Job, and every screen. A starting agent is published before it is ready,
because it is a gossip peer as soon as it starts. Every agent
bootstraps to catalog:8787 and keeps re-resolving it for as long as
it runs.
A Library’s claim
Each Library has one catalog claim, <library>-catalog, and every
Job of the Library runs its agent on it: the walks, the Jobs that
fill gaps, and the cleanup of a deleted Library. The claim keeps the
agent’s actor id and its rows between runs, so each Job syncs only
what changed since the last one.
Two agents must never open one database. ReadWriteOnce limits a
volume to one node, and every pod on that node can still mount it, so
the operator starts a Job of a Library only when no other Job of
that Library is unfinished. On a per-node class the claim is also
ReadWriteOncePod, which limits it to one pod in the whole cluster.
The next Job can run on another node and use that node’s copy. On
every other class the operator writes ReadWriteOnce, and its rule of
one Job at a time is the only guard.
A Job whose pod cannot start holds that rule for the Library until
the Job’s deadline of two hours. The Library’s phase is Blocked
while it waits, and A Job that does not start or that
fails
says what to do.
A claim’s access mode cannot change after it is created, and the
operator creates each claim once. So a <library>-catalog claim made
by an earlier release keeps its mode. To move it to ReadWriteOncePod
on a per-node class, delete the claim after the Library’s last Job
completed and while no other Job of it runs. The next pass creates
the claim again, and the next Job syncs the catalog onto it. A Job
that completed has handed its rows to a catalog pod, so no row is
lost.
Earlier releases gave each Library three more claims:
<library>-enrich-catalog, <library>-trickplay-catalog, and
<library>-trailers-catalog. The operator deletes them, and the
CronJob named <library>-scan, the first time it reconciles the
Library.
How a Job confirms its rows landed
Every worker Job writes a runs row when it starts. It updates that
row when it finishes. The finished write returns the writing agent’s id
and the database version of that write. The Job then records that
agent id and database version in the row, and waits for confirmation.
In a Library’s Job, the close container makes this write after
every phase has ended, so it covers every row the walk and the phases
wrote through the one agent.
Every catalog pod runs a confirmer beside its agent. The confirmer
follows each finished run. It reads crsql_db_versions and
__corro_bookkeeping_gaps from its own copy. Those tables show whether
the copy has every version from that agent through the version named by
the run. Once it does, the confirmer writes a row to confirmations
under its own pod name. The Job exits on the first confirmation for
its run and version. The confirmation means that this catalog copy received
every version from that agent through the version recorded by the Job.
A receiving agent can apply versions out of order. It records missing
earlier versions as gaps, so the newest version does not prove that
earlier versions arrived. While it waits, the Job rewrites its
runs row every ten seconds with a later finish time. Each rewrite is a
new broadcast to current peers. A Job that waits more than two minutes
fails, and Kubernetes retries it. HANDOFF_TIMEOUT sets that limit. The
rows remain safe on the Library’s claim.
A confirmation stays only as long as the run it answers. When a runs
row is deleted, or a later Job of the same worker replaces it, the
confirmer deletes every confirmation whose run is gone, whichever pod
wrote it. It does the same each time its run stream opens, so a delete
made while the confirmer was down is answered when it starts. The
cleanup Job of a deleted Library relies on this. It deletes its own
runs row after the confirmation, and exits only when that
confirmation leaves its copy, because a catalog pod deletes it only
after that pod holds the delete. The same HANDOFF_TIMEOUT bounds this
wait.
Before a phase reads its first gap, it waits until the local copy holds
the newest run a catalog pod confirmed for the Library, the one the
reporter last published. A gap read against a copy that has not synced
would miss titles, or ask a provider again about a title whose attempt
has not arrived. SYNC_TIMEOUT bounds the wait at ten minutes.
How a screen syncs
A screen pod runs the same agent as a native sidecar, and the browser
starts only after the agent’s startup probe passes. The agent’s claim
is named <screen-pod>-catalog, sized from the Catalog, and classed
by spec.screens.storageClassName. A screen pod is pinned to the
machine that holds its display, so a node-local class fits.
A screen holds a second claim beside it, <screen-pod>-art, where
the browser keeps every piece of art it scaled: posters, backdrops,
episode stills, logos, and headshots. spec.screens.artCache.size
sizes it, 2Gi by default, and it takes the same class. The browser
keeps its cache 128 MiB under that size, which is the room an atomic
write needs. Both claims are created once and never updated, so a
size change reaches new screens and not existing ones. To resize an
existing screen, delete its claim, and the next pass creates it at
the new size.
The catalog claim preserves the screen’s database across restarts.
With an emptyDir, a replacement pod must sync the whole catalog.
With a claim, it reuses the database and receives any missing updates.
The following recorded comparison measured restart time and memory
use. These measurements are not performance guarantees:
| Screen restart | On an emptyDir |
On a claim |
|---|---|---|
| Time to the full catalog | 157 s | 0 to 1 s |
| Agent memory after | 211 MiB | 11 MiB |
| First start on a fresh claim | every start | once, 152 s |
A screen in a namespace with no Catalog runs on an emptyDir and
syncs the whole catalog whenever its pod is replaced.
A node-local class binds the claim to the node the pod first landed
on. If the display moves to another machine, the pod cannot schedule
there. After more than five minutes unschedulable, the operator checks
for a bound catalog claim owned by that Player. If it finds one on
a class other than per-node, it deletes the pod and catalog claim,
and removes the art claim if that claim is also bound and owned by
the Player. The next pass recreates them for the new node. This
recovery does not delete claims on a per-node class, since those
claims do not pin the pod to one node.
Reading the catalog by hand
The catalog pod’s images have no shell. The agent’s own binary answers
queries and lists the cluster’s members. For a Catalog named
media, query replica zero:
kubectl -n media exec media-catalog-0 -c catalog -- /corrosion query "SELECT COUNT(*) FROM movies"
kubectl -n media exec media-catalog-0 -c catalog -- /corrosion cluster members
The tables are in the repository at corrosion/schema/catalog.sql,
with a comment on each. Every table is keyed by its library first, so
two libraries in one namespace never share a row.