Troubleshoot

A liken machine reports its problems in the Machine resource. Read in this order: the phase, then the condition that is False, then the status field that the condition’s message names.

kubectl get machines
kubectl describe machine <name>

The phase gives the machine’s state in one word. Ready means every condition is true. Blocked means a change exists that the system refuses to apply, and more time will not fix it. Lost means the machine’s heartbeat stopped. The sections below start from the symptom you see.

The machine stays at the stick menu

The menu has no timeout, so a machine at the menu waits for a person. Select an entry.

If the machine returned to the menu after an installation, the stick is still connected. The stick is first in the boot order. Power the machine off, remove the stick, and power it on again. It then boots from its own disk.

The installation refuses a disk

install as <name> claims blank disks only, so it does not erase data that it did not write. To replace an installation that liken made, select wipe and reinstall as <name> instead. Install a cluster explains both paths, and the disk-replacement case between them.

The installation fails and holds

A failed installation prints its error, lists the machine’s disks, and holds the console. Correct the cause and boot the installation again. An installation is idempotent, so a second attempt is safe.

If the error says that a disk needs a driver the boot path does not have, run the hardware report. The report names the driver and keeps that disk out of the proposed layout. Install a cluster describes the report.

A machine does not join the cluster

Install the first leader first. If a follower boots before any leader serves, the follower waits for a leader.

A declared network port with no cable delays each boot, because the machine waits a maximum of thirty seconds for that port. Remove the port from the manifest, or connect the cable.

If the machine still does not appear in kubectl get machines, watch its console. Every step of the boot prints there, and a boot that fails holds its reason on the screen.

A machine shows Lost

The machine stopped sending its heartbeat, so the cluster operator wrote the phase on its behalf. The cluster operator marks a machine Lost when its heartbeat lease in liken-system has not changed for 40 seconds. It measures the 40 seconds on its own clock, from the moment it saw the last change, so a machine whose clock is wrong does not read Lost for that reason. The MachineLost event gives the time of the last heartbeat, on the cluster operator’s clock. The machine is powered off, or it cannot reach the cluster. When the machine returns, its next report overwrites the phase. If the machine is running and Lost, check its network path to the leaders, and check whether liken-machine-operator restarted on that node. The operator stops its heartbeat and ends its own process when its reconcile loop has been busy for 60 seconds, and the kubelet then starts it again.

A machine shows Blocked

A change exists that the system refuses to stage. The condition that is False names the reason:

A machine returned to the old version

The machine booted the new version one time, the trial failed, and the machine returned to the slot it had proved. Its conditions show RejectedLastBoot, and status.boot.systemRejection records what happened. The machine does not boot that version again until spec.version points at a different release. Roll back describes the fallback and the correction.

The fleet stays on the old version

kubectl get machines shows each machine’s version in the LIKEN column. For a machine that did not move, read its conditions:

If the Cluster’s Progressing condition is False with the reason RolloutStalled, a machine with a granted turn did not return. The cluster grants no more turns until you examine that machine, so the rest of the fleet is safe while you do.

The Cluster’s status.rebootTurns names each machine that holds a reboot turn, and each entry counts against the disruption budget. An entry can name a machine with no RebootApproved condition, because cluster-operator records a turn before it grants it, and a copy that stopped between the two writes leaves the record alone. cluster-operator then writes the grant itself, when the machine is available, and the turn goes on as normal. It removes the entry when the machine holds no grant and its resourceVersion differs from the entry’s machineResourceVersion, or when the Machine is gone. A machine that is down keeps its entry until it returns or reads Lost.

If the Progressing message says that reboot turns wait for a Cluster schema that stores status.rebootTurns, the API server serves the Cluster CRD of an older release, and drops the field from each write. That happens for a short time after the first leader boots a new release, and while a leader on an older release starts k3s again. cluster-operator grants no turn until the newer CRD returns. The leaders apply it again when each one boots the newer release.

A machine crashed

The machine keeps a record of a kernel crash through the reboot. The next boot reads the crash from the machine’s crash journal and reports it in the Machine’s status.lastCrash , with the log lines that the kernel wrote as it failed:

kubectl get machine <name> -o jsonpath='{.status.lastCrash}' | jq

A module parameter did not take

Read back what the kernel holds for the module:

kubectl get machine <name> -o jsonpath='{.status.modules}' | jq

Each declared parameter that the kernel offers a readable file for appears under that module’s parameters, with the value in the kernel’s own spelling. A parameter that is missing there usually has a wrong name. The kernel loads the module anyway and says so only once, as unknown parameter ... ignored, in the machine’s log stream at the moment of the load. If that line is in the log, check the parameter’s spelling against the driver’s documentation and correct the spec. If it is not, the parameter is real but the kernel offers no readable file for it; some drivers register a parameter without one, and the setting still applied at the load.

If the parameter’s name is right but the ModuleParametersApplied condition is False, the condition’s message names one of two cases. The module is built into the kernel, so no load happened and the setting must go elsewhere. Or the module was already loaded before the declared modules ran, so the load that would have applied your string did not run; the message names what loaded it first.

A machine needs a reboot and nothing is staged

A machine converges to its documents, so a machine that already matches them stages nothing and has no reason to reboot. Some faults clear only at boot anyway. A kernel driver that bound the wrong device holds it until the machine restarts, and no edit takes it back.

liken request-reboot asks for that boot:

./liken request-reboot mycluster <name>

The machine waits for its reboot turn, cordons, and drains, the same as a machine applying a staged change. If its rebootPolicy is Manual, it reports RebootPending and waits for liken approve-reboot . The RebootRequestHonored condition reports the progress of the request.

A pod with a device claim stays Pending

Give a workload a device gives the checks, in order.

The recent history of a machine or the cluster

The operators post a Kubernetes Event for each change of a condition, and for each action they take: a machine marked Lost, a reboot turn granted, a spec refused, a kernel crash found. kubectl describe lists the events of the last hour below the status:

kubectl describe machine <name>
kubectl describe cluster <name>

A Machine and a Cluster are in no namespace, so their events are in the default namespace. kubectl events finds them only with -n default or -A:

kubectl events -n default --for machine/<name>

A Warning event needs a person. A Normal event records a change that went as expected. The API server deletes an event one hour after its last repeat, so the conditions and the status fields keep the facts that must last longer.

The event of a condition change has the condition’s reason, for example StagingRejected or RejectedLastBoot, and its message starts with the condition’s name. The other events are these:

Reason Object Type Meaning
MachineJoined Machine Normal The machine’s operator created the Machine from the boot manifest, because the cluster held none.
MachineLost Machine Warning The heartbeat stopped. The message gives the time of the last heartbeat, on the cluster operator’s clock.
KernelCrashed Machine Warning The boot found a new kernel crash record. status.lastCrash holds it.
BootRefused Machine Warning Init refused a boot and powered the machine off. status.lastFailStop holds the reason.
Cordoned Machine Normal The operator cordoned the machine’s Node before a reboot.
Uncordoned Machine Normal The operator returned the Node to the scheduler after the reboot.
BackstopRepaired Machine Warning A pass that no event started changed the machine. The message names each step that wrote. The operator missed a wake, which is a bug in liken: report it with the message. The machine is correct again.
RebootTurnGranted Machine Normal The cluster granted the machine a reboot turn.
RebootTurnReclaimed Machine Normal The machine no longer needs its turn, and the turn returns to the disruption budget.
FluxDeployKeyMinted Cluster Normal The cluster operator made the deploy key. Register its public half at the forge.
FluxEngineSeeded Cluster Normal The flux engine was absent, and the cluster operator created it from the seed that the release carries.

Reading logs

liken stern tails the logs of many pods at once, with the deployment’s credential:

./liken stern mycluster <pod-name-pattern>

When the logs do not explain a problem, the Cluster’s spec.runtime section raises the log level of k3s or containerd, one field each, and says what each level costs.

Reaching a machine’s filesystem

A liken machine has no shell and no SSH, and its own images hold one binary each, so kubectl exec has nothing to run. kubectl debug gives a shell on the node itself, with the host’s process table and the host’s filesystem mounted at /host:

kubectl debug node/<name> -it --profile=sysadmin --image=busybox:1.37

The pod runs privileged in the host’s PID namespace, so ps lists every process on the machine and kill can signal them. Delete the debug pod when you are done; kubectl debug does not remove it.