Procedures

Procedures are the actions you want the equipment to take when a reservation starts and ends, such as unpark the dome, cool the camera, and open the dust cap, and the reverse at the end. Triggers are actions that run once each time a condition changes to the status they name, while their resource is active, such as parking everything when the weather turns unsafe. Automate the equipment shows how to write them.

Each resource lists its procedures in three fields of its spec. Every device kind, the Telescope, and the Observatory have them:

activation, deactivation, and each trigger’s run are lists of actions that run in order. An action is a target state, so it is safe to run twice: the operator reads what the device reports, and sends nothing when the device is there already. Each kind accepts only the actions it supports, and the CRD refuses the others:

Kind Action Default timeout
Dome, Mount state: Parked or state: Unparked 10 min
DustCap state: Open or state: Closed 10 min
FlatPanel state: Lit or state: Dark, the light 10 min
Camera cool: {celsius, within}: write the setpoint and wait until the sensor is within within, 0.5 °C by default 20 min
Camera warm: {celsius, within}: warm the sensor, then switch the cooler off 10 min
every kind, Telescope, Observatory job: {image, command, args, env}: run a container once as a Kubernetes Job 10 min

A warm that does not reach its setpoint by its timeout still switches the cooler off, and notes where the sensor is, because a cooler cannot warm a sensor above the air around it. The setpoint of a camera’s first cool action is also the setpoint that the operator sends again to a camera whose driver restarts, and the Setpoint column of kubectl get cam. The Activation step’s summary names each camera whose driver has a cooler and whose activation does not cool it.

A job is for what the other actions lack, such as a dew heater’s relay or a webhook. The operator creates a batch/v1 Job in the resource’s namespace, owned by the resource, and the action ends when the Job ends: Done when it succeeds, and Failed when it fails. A failed action’s summary gives the exit code of the Job’s container and the last lines, at most 300 bytes, of its termination message, such as Job observatory-lab-activation-a799431dd4 failed: exit code 3: checking the dew heater. The container’s terminationMessagePolicy is FallbackToLogsOnError, so a script that writes no termination message gives the end of its log. When no pod reports how its container ended, such as after activeDeadlineSeconds, the summary gives the Job’s own reason and message. The Job runs its pod once, with no retry, and its activeDeadlineSeconds is the action’s timeout. Kubernetes deletes it an hour after it ends, so its logs stay that long. The container runs as user 1000 with no capabilities and the RuntimeDefault seccomp profile, on a read-only root filesystem with a writable /tmp, so the pod meets the restricted Pod Security level. The operator adds these variables to its environment, after the action’s own env, and its values replace a variable of the same name:

Variable Value
LIKEN_OBSERVATORY the resource’s Observatory, such as lab
LIKEN_TELESCOPE the resource’s Telescope, for a resource of a telescope
LIKEN_RESOURCE the resource, such as Dome/lab
LIKEN_TRIGGER the trigger, such as activation or triggers[0]
INDI_HOST, INDI_PORT the resource’s INDI server, such as lab-observatory.observatory.svc and 7624

The Job’s name holds the resource, the trigger, and a hash of the transition and of the action’s place in the run, such as observatory-lab-activation-a799431dd4. An operator that restarts during a run finds the Job it created, and creates no second one. A run that runs again, such as a failed activation after the retry annotation, deletes the Job of the earlier run and creates a new one.

kubectl get jobs -n observatory -l observatory.liken.sh/role=job

Every action also takes these fields:

A reference with no kind names the resource itself. A kind of Observatory or Telescope with no name names the resource’s own. A reference that names nothing that exists fails the action, and the message names the field, such as after[0]: no Mount north.

The tree orders the runs. Activation runs the Observatory and then its devices, then the Telescope and then its own devices, then the devices of its trains. Deactivation runs the same tiers in reverse. The runs of one tier run in parallel. So the dome unparks before the mount, and the mount parks before the dome, with no after. An after that names a resource of a later tier waits until the action’s timeout, and the action then fails. The observatory’s tiers run when the first reservation in it activates, under a lock, and a second reservation finds those runs Done. They run again at deactivation when the last telescope in the observatory ends.

A Telescope and an Observatory report the condition Active. A telescope is Active from the start of its reservation’s Activation step until the start of its Deactivation step. An observatory is Active from the Activation step of the first reservation in it until the Deactivation step of the last one. The Observatory’s Active message names the reservations that hold telescopes in it now, such as Active for Reservation west-tonight, and a change of holders keeps its lastTransitionTime. Each run records the lastTransitionTime of Active that it answers, in since, so a trigger runs once for each transition.

A resource’s status.procedures holds the last run of each trigger: its trigger, since, state, start and stop times, summary, and each action with its state, times, and summary. The reservation’s Activation and Deactivation steps copy the actions they waited on into status.steps[].actions, each with its resource, so kubectl describe reservation shows the whole activation. A run that an operator restart interrupted resumes from the record, and runs again each action that is not Done. A failed action fails its run and the step, and the retry annotation runs the failed run again, while a run that is Done stays Done. Each run posts an Event on its resource when it starts and when it ends: ProcedureStarted, ProcedureDone, or the Warning ProcedureFailed.

kubectl get dome lab -n observatory -o jsonpath='{.status.procedures}'
kubectl get rsv east-tonight -n observatory \
  -o jsonpath='{.status.steps[?(@.name=="Activation")].actions}'

Triggers

A trigger runs its actions once for each transition of its condition to the status it names, while its resource is active: from the end of its activation run, or from the start of its Telescope’s or Observatory’s activity when it has no activation, until its deactivation begins. when is {kind, name, type, status, for}, with the same rules as a reference in requires. Each run records the condition’s lastTransitionTime in since. A condition that holds when the resource’s activation ends runs the trigger then, so a dome that unparks in bad weather with no requires parks again at once.

for delays the run until the status has held that long, such as for: 20m, and a change of the status before then cancels it, as for does in a Prometheus alert rule. A trigger whose condition names nothing records one Failed run and one Warning.

A run’s record keeps trigger: triggers[0] as its key, and its summary starts with the condition, such as When WeatherStation lab Safe=False: parked Dome lab. The run’s Events name the condition too: Procedure triggers[0] (WeatherStation lab Safe=False) started: state: Parked.

A run that began runs to its end, whatever its condition does meanwhile, so a weather station that flaps, or that reconnects and reports Safe as Unknown for a moment, does not stop a park halfway. Only the resource’s deactivation, or the operator’s stop, ends a run. The deactivation ends it as Skipped with the reason, and a new copy of the operator resumes a run that the stop interrupted.

A resource runs one procedure at a time, its activation and deactivation included, so its device never receives two targets at once. A run that becomes due while another run of the resource goes on is Pending, with a summary such as When WeatherStation lab Safe=True for 20m: waiting for the run of triggers[0] (WeatherStation lab Safe=False) to end. When that run ends, the waiting run begins only if its condition still holds with the same transition time. Otherwise its record is Skipped, with a summary such as When WeatherStation lab Safe=True for 20m: WeatherStation lab Safe is no longer True, so the run did not begin. A waiting run whose resource’s deactivation began does not begin either.

The tree gives no order to a trigger, so after orders the runs of one transition: in a trigger, after waits for the runs of the other resources’ triggers on the same condition and status. A resource with no such trigger, or one that is not active, is not waited for.

kind: Dome
spec:
  triggers:
  - when: {kind: WeatherStation, name: lab, type: Safe, status: "False"}
    run:
    - state: Parked
      after: [{kind: Mount}]
  - when: {kind: WeatherStation, name: lab, type: Safe, for: 20m}
    run: [state: Unparked]

The example observatory states a whole site this way: the dome unparks while the weather station reports Safe, the mounts unpark, the cap opens, and the camera cools to -10 °C. When the weather turns unsafe, the mounts park and then the dome parks, and after 20 minutes of safe weather the dome unparks again. At the end, the flat panel’s light goes off, the cap closes, the camera warms to 5 °C, the mounts park, and the dome parks.

Testing procedures against the simulators

The INDI server images hold indi_getprop and indi_setprop. They run through kubectl exec with no shell, so a person can drive each simulator from your desktop while a reservation holds a telescope. The observatory’s devices run on the pod lab-observatory, and each telescope’s devices on its own pod, such as east-telescope.

The weather simulator decides Safe from its readings. A strong wind turns it unsafe. The simulator publishes its readings each 60 seconds, so the second command asks it to publish them now:

kubectl -n observatory exec lab-observatory -- indi_setprop "Weather Simulator.WEATHER_CONTROL.Wind;Gust=40;30"
kubectl -n observatory exec lab-observatory -- indi_setprop "Weather Simulator.WEATHER_REFRESH.REFRESH=On"

The mounts park, and then the dome parks. A calm wind turns the weather safe again, and the dome unparks after the 20 minutes of the example’s for:

kubectl -n observatory exec lab-observatory -- indi_setprop "Weather Simulator.WEATHER_CONTROL.Wind;Gust=0;0"
kubectl -n observatory exec lab-observatory -- indi_setprop "Weather Simulator.WEATHER_REFRESH.REFRESH=On"

A person parks a mount by hand the same way, and reads a property with indi_getprop:

kubectl -n observatory exec east-telescope -- indi_setprop "Telescope Simulator.TELESCOPE_PARK.PARK=On"
kubectl -n observatory exec lab-observatory -- indi_getprop "Weather Simulator.SAFETY_STATUS.*"

Each resource’s status.procedures holds the last run of each trigger, and its Events give each run’s start and end:

kubectl get dome lab -n observatory -o jsonpath='{.status.procedures}'
kubectl get events -n observatory --field-selector involvedObject.kind=Dome,involvedObject.name=lab

The simulators keep their park state in the pod. A device whose pod starts again comes back unparked, and a dome comes back with its shutter as the simulator starts it. A real driver reads its park state from the hardware.