Metrics

The machine operator and the cluster operator each serve Prometheus metrics on port 9200. Both paths are /metrics. The base needs no Prometheus: the operators serve these ports whether or not anything reads them, and a cluster with no monitoring stack is complete. An owner who runs the prometheus-operator adds the monitoring component, which carries a PodMonitor for each operator and two Grafana dashboards.

Ports

Every process serves /metrics on port 9200. A pod on the cluster network has a port space of its own, so one number serves every process. A pod on the host network shares the node’s port space, so it takes a port nobody else on the host holds: the machine operator holds 9200 there, and the bluetooth operator 9250.

The machine operator also serves /healthz on its port. It answers 500 once the operator’s reconcile loop has been busy for 60 seconds. At the same limit the operator stops renewing its heartbeat and ends its own process, so the kubelet starts it again. /healthz is for a person or a monitor to read, and no probe depends on it.

Metrics

Every liken operator publishes the runtime layer that the Prometheus client library supplies, go_* and process_*, and these:

Component Metric Type What it says
both liken_build_info{component, version} gauge, info the release each component runs
both liken_reconcile_duration_seconds{kind} histogram how long one reconcile pass takes
both liken_reconcile_errors_total{kind} counter passes that failed
both liken_watch_restarts_total{kind} counter watches that closed and opened again

The machine operator and the cluster operator each publish the metrics of their own domain:

Component Metric Type Why
machine-operator liken_release_info{version, slot} gauge, info release per node; fleet skew
machine-operator liken_machine_boot_timestamp_seconds gauge uptime; reboot count over time
machine-operator liken_machine_change_pending{tier} gauge restart, reboot; waiting on approval
machine-operator liken_machine_converged gauge spec matches what runs
machine-operator liken_release_download_bytes_total counter staging progress
machine-operator liken_release_download_failures_total counter a download that keeps failing
machine-operator liken_devices{class} gauge a missing GPU or adapter shows as a drop
machine-operator liken_last_crash_timestamp_seconds gauge pstore capture; a crash loop is a line
machine-operator liken_machine_longest_pass_seconds gauge the longest pass since the operator started; the heartbeat stops at 60
machine-operator liken_machine_backstop_repairs_total{step} counter a write by the five-minute backstop pass, which names a missing wake
machine-operator liken_machine_passes_total{cause} counter passes by what started each one; a settled machine runs few
cluster-operator liken_machines{phase} gauge fleet by phase
cluster-operator liken_disruption_approvals_pending gauge approvals outstanding
cluster-operator liken_machines_behind_target gauge nodes not yet on the target release

The monitoring component

Add this line to the fleet repository’s kustomization, beside the resources it already applies:

components:
  - https://github.com/liken-sh/liken//liken/deploy/monitoring?ref=<version>

Set <version> to the release that the fleet runs.