
# Metrics

The machine operator and the cluster operator each serve Prometheus
metrics on port 9200. Both paths are `/metrics`.
The base needs no Prometheus: the operators serve these ports whether
or not anything reads them, and a cluster with no monitoring stack is
complete. An owner who runs the prometheus-operator adds the
`monitoring` component, which carries a `PodMonitor` for each
operator and two Grafana dashboards.

## Ports

Every process serves `/metrics` on port 9200. A pod on the cluster
network has a port space of its own, so one number serves every
process. A pod on the host network shares the node's port space, so it
takes a port nobody else on the host holds: the machine operator holds
9200 there, and the bluetooth operator 9250.

The machine operator also serves `/healthz` on its port. It answers
`500` once the operator's reconcile loop has been busy for 60 seconds.
At the same limit the operator stops renewing its heartbeat and ends its
own process, so the kubelet starts it again. `/healthz` is for a person
or a monitor to read, and no probe depends on it.

## Metrics

Every liken operator publishes the runtime layer that the Prometheus
client library supplies, `go_*` and `process_*`, and these:

| Component | Metric | Type | What it says |
| --- | --- | --- | --- |
| both | `liken_build_info{component, version}` | gauge, info | the release each component runs |
| both | `liken_reconcile_duration_seconds{kind}` | histogram | how long one reconcile pass takes |
| both | `liken_reconcile_errors_total{kind}` | counter | passes that failed |
| both | `liken_watch_restarts_total{kind}` | counter | watches that closed and opened again |

The machine operator and the cluster operator each publish the
metrics of their own domain:

| Component | Metric | Type | Why |
| --- | --- | --- | --- |
| machine-operator | `liken_release_info{version, slot}` | gauge, info | release per node; fleet skew |
| machine-operator | `liken_machine_boot_timestamp_seconds` | gauge | uptime; reboot count over time |
| machine-operator | `liken_machine_change_pending{tier}` | gauge | `restart`, `reboot`; waiting on approval |
| machine-operator | `liken_machine_converged` | gauge | spec matches what runs |
| machine-operator | `liken_release_download_bytes_total` | counter | staging progress |
| machine-operator | `liken_release_download_failures_total` | counter | a download that keeps failing |
| machine-operator | `liken_devices{class}` | gauge | a missing GPU or adapter shows as a drop |
| machine-operator | `liken_last_crash_timestamp_seconds` | gauge | pstore capture; a crash loop is a line |
| machine-operator | `liken_machine_longest_pass_seconds` | gauge | the longest pass since the operator started; the heartbeat stops at 60 |
| machine-operator | `liken_machine_backstop_repairs_total{step}` | counter | a write by the five-minute backstop pass, which names a missing wake |
| machine-operator | `liken_machine_passes_total{cause}` | counter | passes by what started each one; a settled machine runs few |
| cluster-operator | `liken_machines{phase}` | gauge | fleet by phase |
| cluster-operator | `liken_disruption_approvals_pending` | gauge | approvals outstanding |
| cluster-operator | `liken_machines_behind_target` | gauge | nodes not yet on the target release |

## The monitoring component

Add this line to the fleet repository's kustomization, beside the
resources it already applies:

```yaml
components:
  - https://github.com/liken-sh/liken//liken/deploy/monitoring?ref=<version>
```

Set `<version>` to the release that the fleet runs.

