Roll back
Two mechanisms return a machine to a version that works. If a new version fails its first boot, the machine returns to the previous version without help. You roll the fleet back deliberately when you point the Cluster at an earlier version.
The automatic fallback
Every machine keeps two boot slots, A and B. The machine runs from one slot and writes downloaded releases into the other. An upgrade reboots into the new slot one time only, as a trial:
- If the new kernel panics, the machine resets, and the firmware boots the proven slot. No software is involved.
- If the new version boots but does not rejoin the cluster in ten minutes, a watchdog reboots the machine. The machine then starts again on the proven slot.
In both cases, the machine serves on the version it ran before. Its
phase shows Blocked, its conditions show RejectedLastBoot, and
status.boot.systemRejection
records what happened. The rejection stays until you point
spec.version
at a
different version, so the machine does not boot the failed version
again.
A bad release is never published again. To correct a bad release,
publish a release with the next serial number, add it to the catalog,
and point spec.version at it.
Deliberate rollback
To move the fleet back to an earlier release, point spec.version at
that release:
kubectl edit cluster
The version must still be in
spec.releases.catalog
.
This is a reason to keep the old entries. The rollout is the same as
for an upgrade: each machine downloads the older release into its
inactive slot, verifies it, and reboots on its granted turn, one
machine at a time. The cluster continues to serve during the rollout.
The oldest release you can roll back to
The API refuses a spec.version older than 2026.09.28-001, while
your edit is still open. The cluster operator grants the reboot turns,
and only the copy that holds its leader election Lease acts. A
rollback runs the cluster operator under the RBAC of the target
release, and 2026.09.28-001 is the first release whose RBAC lets it
create and update that Lease. Under an older release’s RBAC, no copy
could take the Lease, so no machine would get its reboot turn and
the rollback would stop.
A Cluster that already names an older version keeps that value. The
API checks the rule only when an edit changes spec.version.
A release from before
status.rebootTurns
runs its rollout without that record. Its cluster-operator writes the
Cluster’s status without the field, so the list is empty while that
release leads, and a copy that paused can then land a grant beyond the
budget, as on every release before the record. When you roll forward,
the newer cluster-operator adds a record for each grant that the
older release wrote, before it grants a new turn.
Before you roll back past a spec field
A release later than 2026.10.08-002 boots a proven manifest that
names fields it does not know. It leaves those fields out, and the
console names each one, for example the proven manifest names fields this release does not know, and this boot leaves them out: spec.serio. The machine keeps its storage, network, and every other
field. The operator of that release also drops the conditions that
only a newer release writes, such as SerioAttached, so a stale
condition does not hold the machine Degraded. You can roll back to
such a release without the steps below.
A release up to and including 2026.10.08-002 cannot read a manifest
that uses a field newer than the release. spec.serio is such a
field: releases before it reject a manifest that declares it. A
machine that declares spec.serio and boots such a release cannot
use its proven manifest, and it falls back to the seed manifest in
the release’s image on that slot:
- If the image carries a seed for the machine, the machine boots under it. The seed is the install-time manifest, so its storage and network can be stale, and the boot records it as the new proven manifest.
- If the image carries no seed, the machine boots under an empty
Machine. - If the seed declares
spec.seriotoo, the older release cannot read it either, and the machine powers off. Recovery needs a new install stick.
To roll back such a machine to a release up to and including
2026.10.08-002, remove the field first:
-
Remove
spec.seriofrom theMachine:kubectl patch machine <name> --type=json \ -p '[{"op":"remove","path":"/spec/serio"}]' -
Wait until the
SpecConvergedcondition shows that the removal is staged: the reasonRebootPendingorAwaitingTurn. WithrebootPolicy: Auto, the reason isRebootRequested, and the machine reboots at once. Then wait until the condition isTrueafter that reboot. A removed entry stays attached until the boot that applies the removal. -
Point
spec.versionat the older release. The older release boots the staged manifest, or the proven one that the reboot in step 2 promoted, and it can read both.
The older release does not know the SerioAttached condition, so it
keeps whatever value the condition had. If that value is False, the
machine reads as Degraded after the rollback. Remove the condition
from the status with kubectl edit machine <name> --subresource=status.