Roll back

Two mechanisms return a machine to a version that works. If a new version fails its first boot, the machine returns to the previous version without help. You roll the fleet back deliberately when you point the Cluster at an earlier version.

The automatic fallback

Every machine keeps two boot slots, A and B. The machine runs from one slot and writes downloaded releases into the other. An upgrade reboots into the new slot one time only, as a trial:

In both cases, the machine serves on the version it ran before. Its phase shows Blocked, its conditions show RejectedLastBoot, and status.boot.systemRejection records what happened. The rejection stays until you point spec.version at a different version, so the machine does not boot the failed version again.

A bad release is never published again. To correct a bad release, publish a release with the next serial number, add it to the catalog, and point spec.version at it.

Deliberate rollback

To move the fleet back to an earlier release, point spec.version at that release:

kubectl edit cluster

The version must still be in spec.releases.catalog . This is a reason to keep the old entries. The rollout is the same as for an upgrade: each machine downloads the older release into its inactive slot, verifies it, and reboots on its granted turn, one machine at a time. The cluster continues to serve during the rollout.

The oldest release you can roll back to

The API refuses a spec.version older than 2026.09.28-001, while your edit is still open. The cluster operator grants the reboot turns, and only the copy that holds its leader election Lease acts. A rollback runs the cluster operator under the RBAC of the target release, and 2026.09.28-001 is the first release whose RBAC lets it create and update that Lease. Under an older release’s RBAC, no copy could take the Lease, so no machine would get its reboot turn and the rollback would stop.

A Cluster that already names an older version keeps that value. The API checks the rule only when an edit changes spec.version.

A release from before status.rebootTurns runs its rollout without that record. Its cluster-operator writes the Cluster’s status without the field, so the list is empty while that release leads, and a copy that paused can then land a grant beyond the budget, as on every release before the record. When you roll forward, the newer cluster-operator adds a record for each grant that the older release wrote, before it grants a new turn.

Before you roll back past a spec field

A release later than 2026.10.08-002 boots a proven manifest that names fields it does not know. It leaves those fields out, and the console names each one, for example the proven manifest names fields this release does not know, and this boot leaves them out: spec.serio. The machine keeps its storage, network, and every other field. The operator of that release also drops the conditions that only a newer release writes, such as SerioAttached, so a stale condition does not hold the machine Degraded. You can roll back to such a release without the steps below.

A release up to and including 2026.10.08-002 cannot read a manifest that uses a field newer than the release. spec.serio is such a field: releases before it reject a manifest that declares it. A machine that declares spec.serio and boots such a release cannot use its proven manifest, and it falls back to the seed manifest in the release’s image on that slot:

To roll back such a machine to a release up to and including 2026.10.08-002, remove the field first:

  1. Remove spec.serio from the Machine:

    kubectl patch machine <name> --type=json \
      -p '[{"op":"remove","path":"/spec/serio"}]'
    
  2. Wait until the SpecConverged condition shows that the removal is staged: the reason RebootPending or AwaitingTurn. With rebootPolicy: Auto, the reason is RebootRequested, and the machine reboots at once. Then wait until the condition is True after that reboot. A removed entry stays attached until the boot that applies the removal.

  3. Point spec.version at the older release. The older release boots the staged manifest, or the proven one that the reboot in step 2 promoted, and it can read both.

The older release does not know the SerioAttached condition, so it keeps whatever value the condition had. If that value is False, the machine reads as Degraded after the rollback. Remove the condition from the status with kubectl edit machine <name> --subresource=status.