Upgrade the fleet

One edit to the Cluster moves every machine to a new release. Each machine downloads the release, verifies every artifact, and reboots when the cluster grants its turn. You do not rebuild the media, and you do not touch the machines.

1. Find the release

The channel at releases.liken.sh lists every release. The cluster also polls the channel:

kubectl get clusters

The AVAILABLE column shows the latest version that the channel announces. To make the cluster poll the channel now, set the liken.sh/check-releases annotation to a new value:

kubectl annotate cluster --all --overwrite liken.sh/check-releases="$(date -Is)"

The value of the annotation has no meaning. The change of the value is the request.

An upgrade needs two facts: the version, and the digest of that release’s release.yaml. The release’s page on GitHub gives both, as a catalog entry you can copy. To compute the digest yourself:

curl -fsSL https://releases.liken.sh/<version>/release.yaml | sha256sum

2. Catalog the release and set the target

kubectl edit cluster

Add the release to spec.releases.catalog , and point spec.version at it:

spec:
  version: "2026.10.10-001"
  releases:
    source: https://releases.liken.sh
    catalog:
      - version: "2026.10.10-001"
        digest: sha256:<hex>

If spec.version names no catalog entry, the API refuses the change while your edit is still open. The digest is the start of the trust chain . The digest names the release document, and the release document names the artifacts. Each machine checks every downloaded byte against one or the other.

3. The rollout

Each machine that runs a different version:

  1. Downloads the release into the boot slot it does not run from, and verifies each artifact against the digest chain.
  2. Stages the change and asks the cluster for a reboot turn.
  3. Cordons and drains its node when the cluster grants the turn. The PodDisruptionBudgets of the workloads apply during the drain. The drain removes a pod that holds a DRA claim before the pod that runs its driver. The driver stays to answer the kubelet’s unprepare call.
  4. Reboots into the new slot one time, as a trial. The trial is a success when the OS starts and rejoins the cluster. From that time, the machine boots that slot. If the trial fails, the machine returns to the other slot without help: Roll back describes this.

The cluster grants turns within spec.disruption.maxUnavailable (the default is one machine at a time). Only one leader is down at a time, whatever the budget says, because the datastore needs a majority of the leaders.

cluster-operator records each turn in the Cluster’s status.rebootTurns before it grants the turn, and the record counts against the budget. The record is a write that names the Cluster’s resourceVersion, so two copies of cluster-operator cannot both give one slot of the budget to a machine. A copy that paused, for example on a stalled node, can send a grant after another copy took over. That grant meets the record and adds no turn. A machine’s record leaves the list after the cluster takes its grant back.

When cluster-operator starts, for example after the machine that ran it loses power, it grants no turn for its first 40 seconds. A machine that went down before cluster-operator started still shows its last status until 40 seconds pass with no heartbeat. After that it reads Lost and counts against the budget. The Cluster’s Progressing message gives the time that granting resumes.

If a machine’s rebootPolicy is Manual (the default), the machine stages the change, reports RebootPending, and waits for you. Grant the reboot with liken approve-reboot :

liken approve-reboot mycluster <machine>

The machine then takes its turn under the same disruption budget as an Auto machine: it drains first, and only one leader is ever down at a time. Set rebootPolicy: Auto on machines that must take their turn without an operator.

You can change spec.version again before every machine reboots. A machine that staged the earlier target withdraws that release first, stops its download if one runs, and then downloads the new target into the same slot. A reboot in the meantime tries neither release: the earlier one is no longer staged, and the new one is staged only after its download verifies. Before each trial, the machine checks the slot’s release document and every artifact against the staged release, and it reboots without the trial when they do not match.

4. Watch the rollout

kubectl get machines

The LIKEN column changes to the new version one machine at a time, and the STATUS column shows each machine’s step in the rollout:

The phase of the Cluster shows Updating during the rollout, and Ready when the rollout is complete.

The TURNS column of the wide listing names each machine that holds a reboot turn:

kubectl get clusters -o wide

NAME   STATUS     MACHINES   VERSION          NEWEST           AVAILABLE        LEADERS                        ENDPOINT   AGE   TURNS
lab    Updating   4/5        2026.10.12-001   2026.10.12-001   2026.10.12-001   ["node-1","node-2","node-3"]   ...        9d    node-4

If a machine with a granted turn does not return, the cluster sets its Progressing condition to False with the reason RolloutStalled. The cluster grants no more turns until you examine the machine.