Troubleshoot
A liken machine reports its problems in the Machine resource. Read
in this order: the phase, then the condition that is False, then
the status field that the condition’s message names.
kubectl get machines
kubectl describe machine <name>
The phase gives the machine’s state in one word. Ready means every
condition is true. Blocked means a change exists that the system
refuses to apply, and more time will not fix it. Lost means the
machine’s heartbeat stopped. The sections below start from the
symptom you see.
The machine stays at the stick menu
The menu has no timeout, so a machine at the menu waits for a person. Select an entry.
If the machine returned to the menu after an installation, the stick is still connected. The stick is first in the boot order. Power the machine off, remove the stick, and power it on again. It then boots from its own disk.
The installation refuses a disk
install as <name> claims blank disks only, so it does not erase
data that it did not write. To replace an installation that liken
made, select wipe and reinstall as <name> instead.
Install a cluster
explains both paths, and the disk-replacement case between them.
The installation fails and holds
A failed installation prints its error, lists the machine’s disks, and holds the console. Correct the cause and boot the installation again. An installation is idempotent, so a second attempt is safe.
If the error says that a disk needs a driver the boot path does not have, run the hardware report. The report names the driver and keeps that disk out of the proposed layout. Install a cluster describes the report.
A machine does not join the cluster
Install the first leader first. If a follower boots before any leader serves, the follower waits for a leader.
A declared network port with no cable delays each boot, because the machine waits a maximum of thirty seconds for that port. Remove the port from the manifest, or connect the cable.
If the machine still does not appear in kubectl get machines,
watch its console. Every step of the boot prints there, and a boot
that fails holds its reason on the screen.
A machine shows Lost
The machine stopped sending its heartbeat, so the cluster operator
wrote the phase on its behalf. The cluster operator marks a machine
Lost when its heartbeat lease in liken-system has not changed for
40 seconds. It measures the 40 seconds on its own clock, from the
moment it saw the last change, so a machine whose clock is wrong does
not read Lost for that reason. The MachineLost event gives the
time of the last heartbeat, on the cluster operator’s clock. The
machine is powered off, or it cannot reach the cluster. When the
machine returns, its next report overwrites the phase. If the machine
is running and Lost, check its network path to the leaders, and check
whether liken-machine-operator restarted on that node. The operator
stops its heartbeat and ends its own process when its reconcile loop has
been busy for 60 seconds, and the kubelet then starts it again.
A machine shows Blocked
A change exists that the system refuses to stage. The condition that
is False names the reason:
StagingRejected: the spec asks for something the machine cannot do. The common case is a storage role that is smaller in the spec than on the disk, after a reinstallation with a different layout. The condition’s message gives the sizes. Install a cluster gives the order of the corrections.RejectedLastBoot: the machine booted the change one time, and the boot failed. See the next section.
A machine returned to the old version
The machine booted the new version one time, the trial failed, and
the machine returned to the slot it had proved. Its conditions show
RejectedLastBoot, and
status.boot.systemRejection
records what happened. The machine does not boot that version again
until spec.version
points at a different release. Roll back
describes the
fallback and the correction.
The fleet stays on the old version
kubectl get machines shows each machine’s version in the LIKEN
column. For a machine that did not move, read its conditions:
RebootPending: the machine’srebootPolicyisManual, which is the default, and the machine waits for you. Read what it waits for and grant the reboot withliken approve-reboot.AwaitingTurn: the machine waits for the cluster’s disruption budget. The default budget is one machine at a time, and only one leader is down at a time whatever the budget says.Downloading: the machine still fetches or verifies the release. A slow link makes this step long. A download that receives no bytes for one minute stops, and the machine retries it on its own. A download that fails waits 10 seconds before the first retry, and the wait doubles after each failure up to 2 minutes. A change of the target release starts the new download without that wait. The condition’s message gives the reason of the last failure and the time of the next retry.StagingFailed: the machine could not write or withdraw a staged record. The message gives the error. When the target changed while another release was staged, the machine does not download the new release until it withdraws the old record.AwaitingPodRefresh: the machine runs the new release and waits for the pod steward to replace its operator pod from the new template. The steward does that after a leader boots the release.
If the Cluster’s Progressing condition is False with the reason
RolloutStalled, a machine with a granted turn did not return. The
cluster grants no more turns until you examine that machine, so the
rest of the fleet is safe while you do.
The Cluster’s
status.rebootTurns
names each machine that holds a reboot turn, and each entry counts
against the disruption budget. An entry can name a machine with no
RebootApproved condition, because cluster-operator records a turn
before it grants it, and a copy that stopped between the two writes
leaves the record alone. cluster-operator then writes the grant
itself, when the machine is available, and the turn goes on as
normal. It removes the entry when the machine holds no grant and its
resourceVersion differs from the entry’s machineResourceVersion,
or when the Machine is gone. A machine that is down keeps its entry
until it returns or reads Lost.
If the Progressing message says that reboot turns wait for a Cluster
schema that stores status.rebootTurns, the API server serves the
Cluster CRD of an older release, and drops the field from each write.
That happens for a short time after the first leader boots a new
release, and while a leader on an older release starts k3s again.
cluster-operator grants no turn until the newer CRD returns. The
leaders apply it again when each one boots the newer release.
A machine crashed
The machine keeps a record of a kernel crash through the reboot. The
next boot reads the crash from the machine’s crash journal and reports
it in the Machine’s
status.lastCrash
, with
the log lines that the kernel wrote as it failed:
kubectl get machine <name> -o jsonpath='{.status.lastCrash}' | jq
A module parameter did not take
Read back what the kernel holds for the module:
kubectl get machine <name> -o jsonpath='{.status.modules}' | jq
Each declared parameter that the kernel offers a readable file for
appears under that module’s parameters, with the value in the
kernel’s own spelling. A parameter that is missing there usually has
a wrong name. The kernel loads the module anyway and says so only
once, as unknown parameter ... ignored, in the machine’s log
stream at the moment of the load. If that line is in the log, check
the parameter’s spelling against the driver’s documentation and
correct the spec. If it is not, the parameter is real but the
kernel offers no readable file for it; some drivers register a
parameter without one, and the setting still applied at the load.
If the parameter’s name is right but the ModuleParametersApplied
condition is False, the condition’s message names one of two
cases. The module is built into the kernel, so no load happened and
the setting must go elsewhere. Or the module was already loaded
before the declared modules ran, so the load that would have applied
your string did not run; the message names what loaded it first.
A machine needs a reboot and nothing is staged
A machine converges to its documents, so a machine that already matches them stages nothing and has no reason to reboot. Some faults clear only at boot anyway. A kernel driver that bound the wrong device holds it until the machine restarts, and no edit takes it back.
liken request-reboot
asks for that boot:
./liken request-reboot mycluster <name>
The machine waits for its reboot turn, cordons, and drains, the same
as a machine applying a staged change. If its rebootPolicy is
Manual, it reports RebootPending and waits for
liken approve-reboot
.
The RebootRequestHonored condition reports the progress of the
request.
A pod with a device claim stays Pending
Give a workload a device gives the checks, in order.
The recent history of a machine or the cluster
The operators post a Kubernetes Event for each change of a
condition, and for each action they take: a machine marked Lost, a
reboot turn granted, a spec refused, a kernel crash found.
kubectl describe lists the events of the last hour below the
status:
kubectl describe machine <name>
kubectl describe cluster <name>
A Machine and a Cluster are in no namespace, so their events are
in the default namespace. kubectl events finds them only with
-n default or -A:
kubectl events -n default --for machine/<name>
A Warning event needs a person. A Normal event records a change
that went as expected. The API server deletes an event one hour after
its last repeat, so the conditions and the status fields keep the
facts that must last longer.
The event of a condition change has the condition’s reason, for
example StagingRejected or RejectedLastBoot, and its message
starts with the condition’s name. The other events are these:
| Reason | Object | Type | Meaning |
|---|---|---|---|
MachineJoined |
Machine |
Normal |
The machine’s operator created the Machine from the boot manifest, because the cluster held none. |
MachineLost |
Machine |
Warning |
The heartbeat stopped. The message gives the time of the last heartbeat, on the cluster operator’s clock. |
KernelCrashed |
Machine |
Warning |
The boot found a new kernel crash record. status.lastCrash
holds it. |
BootRefused |
Machine |
Warning |
Init refused a boot and powered the machine off. status.lastFailStop
holds the reason. |
Cordoned |
Machine |
Normal |
The operator cordoned the machine’s Node before a reboot. |
Uncordoned |
Machine |
Normal |
The operator returned the Node to the scheduler after the reboot. |
BackstopRepaired |
Machine |
Warning |
A pass that no event started changed the machine. The message names each step that wrote. The operator missed a wake, which is a bug in liken: report it with the message. The machine is correct again. |
RebootTurnGranted |
Machine |
Normal |
The cluster granted the machine a reboot turn. |
RebootTurnReclaimed |
Machine |
Normal |
The machine no longer needs its turn, and the turn returns to the disruption budget. |
FluxDeployKeyMinted |
Cluster |
Normal |
The cluster operator made the deploy key. Register its public half at the forge. |
FluxEngineSeeded |
Cluster |
Normal |
The flux engine was absent, and the cluster operator created it from the seed that the release carries. |
Reading logs
liken stern
tails the logs of
many pods at once, with the deployment’s credential:
./liken stern mycluster <pod-name-pattern>
When the logs do not explain a problem, the Cluster’s
spec.runtime
section raises
the log level of k3s or containerd, one field each, and says what
each level costs.
Reaching a machine’s filesystem
A liken machine has no shell and no SSH, and its own images hold
one binary each, so kubectl exec has nothing to run. kubectl debug
gives a shell on the node itself, with the host’s process table and
the host’s filesystem mounted at /host:
kubectl debug node/<name> -it --profile=sysadmin --image=busybox:1.37
The pod runs privileged in the host’s PID namespace, so ps lists
every process on the machine and kill can signal them. Delete the
debug pod when you are done; kubectl debug does not remove it.