Devices
liken gives hardware to workloads with dynamic resource allocation,
the Kubernetes API for devices. A pod asks for a device by its
properties. The scheduler selects a node that has such a device, and
the node gives that pod the device’s /dev entries. The pod needs no
privilege, no host mounts, and no knowledge of the machine it runs
on.
Give a workload a device
gives the steps.
This page describes what liken publishes, and why.
The driver
The driver’s name is liken.sh. It is not a separate program. It is
two jobs in the machine operator, which runs on every node. One job
reads sysfs and publishes what it finds. The other job answers the
kubelet when a pod’s claim comes to the node.
Usually a cluster administrator installs a DRA driver as a DaemonSet.
liken’s driver is part of the operating system, because the operating
system is the only thing that identifies the hardware before other
software starts.
The device operators are separate DRA drivers that you install as workloads. Each one publishes a kind of device this driver does not, such as the controllers paired to a Bluetooth radio.
What a node publishes
Each node publishes one ResourceSlice, with the name
<node>-liken.sh. The slice lists the node’s claimable devices:
kubectl get resourceslices
kubectl get resourceslice <node>-liken.sh -o yaml
The node considers every device on the PCI bus and the USB bus, and every device on the board that owns a device node. A board device is one that the firmware describes and no bus enumerates: a firmware TPM, a laptop’s own keyboard behind the i8042 controller, the ACPI power button, or a serial port on the board. A device appears in the slice when these three conditions are true:
- A driver is bound to it. Hardware with no driver can supply
nothing, so it goes in the machine’s unclaimed report instead. To
move a device from one report to the other, declare its module in
the Machine’s
spec.modules. A USB device that no driver binds at all is the exception, because a program drives it through libusb. See USB devices with no kernel driver . - A claim on it supplies something. The device’s sysfs subtree
must contain device nodes. A network card is hardware, but it has
nothing to give a pod, so it never appears. A Bluetooth adapter is
the one exception. Its driver,
btusb, puts it in the slice, because the kernel gives a radio no device node at all. See Bluetooth adapters . - The machine does not depend on it. A device whose subtree holds a
storage role belongs to the machine, and never to a workload. A
disk belongs to the machine, through a storage role, or to the
workloads, through a claim, but never to both.
machine-operatorreads which partitions back a role from the facts that init publishes. When it can’t read the facts, it can’t tell a system disk from a spare one, so it offers no disk at all, and the Machine’sFactsPublishedcondition isFalse. A device with no disk in its subtree, such as a GPU, stays in the slice. A claim that already holds a disk fails to start its next container until the facts read again, and the operator checks each disk again when it prepares a claim, so an allocation from an older slice can’t deliver one.
The machine holds four kinds of node, and no claim receives them.
The console, which the kernel lists in
/sys/class/tty/console/active, carries the boot and the kernel’s
messages. /dev/rtc0 is the clock that init writes the system time
into, and the kernel lets one process at a time open it. The TPM’s
nodes, /dev/tpm* and /dev/tpmrm*, give any process that opens them
the TPM itself. On most boards the TPM’s owner and lockout passwords
are empty, so a pod with no privilege could set them, clear the TPM,
or extend its PCRs until the next boot, and each of those reaches every
later user of the TPM. A device whose only nodes these are does not
appear. The fourth is the tty of a
serial line that a spec.serio entry matches. The machine holds that line
attached to the kernel’s serio layer, and a pod that received the tty
could end the attachment under every other claim. See
Serial-line adapters
.
The slice is an allocation offer, not a full record of the machine’s hardware. The scheduler can allocate only what a slice lists, so the slice contents control which devices workloads can claim.
The slice does not list the bus structure. Hubs, PCIe ports, and the USB core’s own devices are the structure that peripherals attach to, not peripherals.
The node walks its devices again when the kernel sends a uevent for a device that arrives, leaves, or changes its driver. The node waits for one second with no further uevent, so that the dozen uevents of one USB device make one walk, and then writes the slice. So a device reaches the slice about a second after the kernel reports it. The uevents for the virtual network devices that pods add and remove wake no walk.
Device names
A device’s name is its bus and its address, in lowercase, with dashes
in place of the punctuation: pci-0000-00-02-0, usb-2-1-1-0. A
board device’s address is the name the firmware gives it, so a
firmware TPM is platform-msft0101-00. A name stops at 63
characters, the limit of a DNS label.
The address gives the slot, not the unit. If you replace a dongle
with an identical dongle in the same port, the device name does not
change, which is what a claim on the adapter on that port needs. To
select one physical unit instead, select on its serial attribute.
A constraint pairs the requests of a claim by an attribute, and never
by name. The address is also an attribute, address, so a claim can
constrain two requests to one physical device.
Attributes
Every attribute except the last two belongs to the driver’s domain, so a
selector reads it as device.attributes["liken.sh"].<name>. If the
hardware does not have an attribute, the attribute is absent, not
empty. Thus has(device.attributes["liken.sh"].serial) gives a
correct result.
| Attribute | Type | What it is |
|---|---|---|
bus |
string | pci or usb, or for a board device, the bus the kernel registered it on, such as platform |
address |
string | the device’s address on that bus: 0000:00:02.0 on PCI, the port path on USB, and the firmware’s name for a board device, such as MSFT0101:00. Every device liken publishes for one physical device has the same address |
driver |
string | the name of the bound driver, such as i915 |
class |
string | the type of device, in one word: display, multimedia, serial-bus |
classCode |
string | the full class code that the bus published: six hex digits on PCI, two on USB |
subsystem |
string | the kind of device that liken published: drm, sound, tty, cec. It is absent when the delivery is a mix liken has no name for, and when the device delivers no node of its own, as a Bluetooth adapter does |
renderNode |
bool | the device supplies a DRM render node |
displayNode |
bool | the device supplies a DRM card node, which provides modesetting |
name |
string | the name of the device in words, from its own strings or from the PCI database |
modalias |
string | the identifier that the kernel uses to match drivers |
serial |
string | the serial number of the hardware, when it has one |
vendor |
string | the vendor ID, lowercase hex with no prefix |
product |
string | the product ID, lowercase hex with no prefix |
sound.liken.sh/supportsSound |
bool | a sound server can run against this device. The attribute carries its own domain, so a selector reads it as device.attributes["sound.liken.sh"].supportsSound. The domain belongs to no single driver, so a DeviceClass that selects the attribute names no driver, and any driver may stamp it on a device that supports a sound server |
resource.kubernetes.io/pciBusID |
string | the address of a PCI device, the same value as address, on every device liken publishes for a PCI device. Kubernetes defines the attribute and its format, Domain:Bus:Device.Function, in the standard device attributes
. The domain belongs to no single driver, so another driver can publish the same attribute for the same card, and a matchAttribute constraint can pair a device of liken with a device of that driver. A USB device has no such attribute |
Use the attribute that describes what you need. renderNode and
classCode describe a capability of the hardware, and they stay
correct across a fleet of different machines. vendor and product
give one model, and a DeviceClass that uses them stops working on the
next machine that you buy.
address is the attribute that pairs devices. Two cards in one
machine have different addresses, and every device liken publishes for
one card has that card’s address. So a claim with a request for a
render node and a request for a card node constrains the two with
matchAttribute: liken.sh/address, and both requests get halves of
the same card. The
guide
has the claim.
liken.sh/address pairs only the devices that liken publishes,
because a bare attribute name belongs to the driver that published it.
To pair a liken device with a device of another driver for the same
card, constrain the requests with
matchAttribute: resource.kubernetes.io/pciBusID.
liken publishes no capability facts that need a driver stack to
measure, for example the codecs that a GPU can encode. To read those
facts, you must run libva and a vendor driver, which the image does
not contain. A pod that holds a claim on the render node can measure
them for itself, so the operating system does not publish them.
media-operator
measures them, and publishes them as devices of its
own driver, media.liken.sh, that pair with the render node through
resource.kubernetes.io/pciBusID.
Sharing
liken allocates a device to one claim, unless liken publishes
allowMultipleAllocations for that device. A device’s subtree can
deliver nodes of more than one kernel subsystem. When it does, liken
can publish more than one device for it, one for each kind of node. A
GPU splits again, because its render node and its card node give
different authority over the same silicon. The name of each extra
device is the primary device’s name plus a suffix. A claim receives
the nodes of the one published device it allocated, and no others.
liken publishes allowMultipleAllocations for one kind of device:
the graphics half of a GPU. A graphics device is one that delivers a
DRM render node, and the published graphics device delivers that
render node, /dev/dri/renderD*. It has renderNode: true. The
driver arbitrates between concurrent clients on a render node, so
more than one claim can hold it.
An audio controller is a device that delivers ALSA’s nodes, and no
nodes except the ones a sound card holds. It has
subsystem: sound and sound.liken.sh/supportsSound: true, and it
is exclusive. In practice one sound server
owns every PCM on a card and mixes its clients’ streams through them,
so the card belongs to one claim. A second claimant stays pending in the
scheduler, and Kubernetes reports that state instead of returning ALSA’s
EBUSY at play time.
A claim on an audio controller delivers the card’s whole subtree,
which is more than the /dev/snd nodes. ALSA registers an input
device for each jack that it can sense, and an HDA controller with
HDMI outputs has one for each display pin, so the claim also delivers
those /dev/input/event* nodes. A jack reports the state of an
output that the same claim plays through, so it is not a separate
device.
A GPU publishes its card node, /dev/dri/card*, as a separate
device. The name of that device is the primary name plus -display,
and it has displayNode: true. This device is exclusive, because
DRM master is one for each card: the kernel gives modesetting to one
open card node, and a second display program on the same card fails
when it starts. An exclusive device makes the second claim stay pending
in the scheduler, where Kubernetes reports its pending state. A workload
that modesets and also renders claims both devices, with one request
for each. On a machine with two cards, a constraint of
matchAttribute: liken.sh/address on those two requests keeps both on
one card.
A GPU driver also registers i2c monitor-control buses. These publish
as their own device, with subsystem: i2c-dev. That device stays
exclusive. An i2c node passes raw transfers to every device on its
wire, and two writers on one wire have no arbitration contract. A
DisplayPort output also registers a DisplayPort AUX channel, with
subsystem: drm_dp_aux_dev. This node publishes the same way, as its
own exclusive device, and it exists only while a display is
connected. The legacy framebuffer node is not delivered at all:
holding it grants display takeover, and no workload claims a bare
framebuffer.
A USB-CEC adapter that a spec.serio entry attaches publishes two
devices, both exclusive: the CEC bus, with subsystem: cec, and the
TV remote’s input device, with the suffix -input and
subsystem: input.
Serial-line adapters
gives the reasons.
A device that delivers one kind of node publishes as one device,
unless it is a GPU. When a device delivers a mix that liken has no
name for, liken publishes the whole device as exclusive and omits
subsystem. A Bluetooth adapter is also published whole and exclusive
with no subsystem, because its only delivered node is its usbfs node.
A DRM render node has a multiplexing contract in the kernel: the driver arbitrates between concurrent clients. A drill measured this on an integrated GPU. Twelve concurrent encoders divided the GPU equally, and two pods each got approximately half of the throughput that one pod got alone. A serial port has no such contract, and a dongle’s control endpoint has none.
Only the driver can state this, because only the driver writes a
ResourceSlice. Thus the rule is narrow. If a device is incorrectly
marked as shareable, two workloads get the same hardware while each
one operates as the only user, and no DeviceClass, claim, or workload
can correct it. If a device is incorrectly marked as exclusive, a
claim stays pending in the scheduler, where Kubernetes reports its
pending state.
Sharing a claim is not the same thing
allowMultipleAllocations controls how many claims can allocate
a device. It does not control how many pods can use one claim.
Every pod that names a ResourceClaim shares that claim, and all of
these pods receive the device. Kubernetes does this on purpose: it is
how two pods share one allocation. liken does not refuse the second
pod. A refusal would be a race with no owner. Also, the delivery of a
device node was never the mechanism that gave exclusive access. The
kernel gives exclusive access, through O_EXCL and the driver’s own
open path.
To give a device to one pod only, use a ResourceClaimTemplate,
which makes a separate claim for each pod. If you use one
ResourceClaim for more than one pod, you share the device on
purpose.
What a claim delivers
A claim delivers device nodes only. The node writes a CDI
specification that names the /dev entries for that claim, and the
container runtime injects them into the containers that requested the
claim. The claim grants no privilege, mounts no host path, and loads
no kernel module for the pod.
The workload supplies everything else that it needs: the userspace library that communicates with the device, and the group membership that its image gives its user.
USB devices
A claim on a USB device also delivers that device’s usbfs node,
/dev/bus/usb/<busnum>/<devnum>. The devices of an attached serial
line are the exception, and
Serial-line adapters
says why. A program that uses libusb, for
example Network UPS Tools, reads sysfs to find the hardware and then
opens this node to communicate with it. A node that a kernel driver
registers, such as hidraw, supplies that driver’s protocol only, so
it cannot take the place of the usbfs node.
A program that uses libusb cannot share an interface with a kernel
driver, so it detaches the kernel driver while it runs. liken
publishes only devices that have a driver, so the device leaves the
node’s slice for as long as the pod runs. When the pod stops and its
claim ends, the node binds a kernel driver to the interface again, and
the device returns to the slice at the next reconcile pass.
The kernel gives the device a new device number at each enumeration. The node changes when you unplug the device and plug it in again. Each reconcile pass writes the current node into the specification of every claim that the kubelet prepared, so the next pod receives the current node. A container that already runs keeps the node that it received, and it receives the current node when the pod restarts.
While the hardware is unplugged, the pass writes a node that does not
exist into the claim’s specification instead. The old node would be
unsafe, because the kernel gives its name to the next device that
arrives: another USB disk becomes /dev/sda. A container that starts
while the hardware is away fails to start, and the kubelet retries it
under its restart backoff. When the hardware returns to the same
port, the next pass writes its nodes back, and the container starts at
its next retry.
usbfs has no interface boundary. If a device has more than one interface with a driver, each interface publishes as its own device, and each one delivers the same usbfs node. A pod that holds one of these claims can communicate with the whole device.
USB devices with no kernel driver
Some USB devices have no kernel driver, because the vendor’s driver is
a program that uses libusb. ZWO’s astronomy cameras, smart card
readers that pcscd serves, and many software-defined radios are this
kind. When no interface of a USB device has a driver, the device
publishes whole, with the name of its port path, such as usb-1-2.
A claim on it delivers the device’s usbfs node, and nothing else. The
class and classCode attributes come from the device’s first
interface, so a DeviceClass can select a smart card reader by
classCode == "0b". The device is exclusive.
A device with a driver on any interface does not publish whole. Its
driven interfaces publish as their own devices, each with the usbfs
node, and a claim on the whole device would hand the same hardware to
a second workload. A device leaves the slice while a program holds one
of its interfaces, because the kernel then shows usbfs as that
interface’s driver, and it returns when the program lets go.
A device whose kernel module is not loaded yet also publishes whole, until the module binds an interface. Then the driven interface publishes in its place.
Bluetooth adapters
A claim on a Bluetooth adapter delivers the adapter’s usbfs node,
/dev/uinput, the 32 legacy input event nodes, and, on a machine
whose spec.modules declares uhid, the /dev/uhid node. The
adapter has no node of its own. A program reaches a radio
through an AF_BLUETOOTH socket, and it binds that socket to an
adapter by index. The kernel registers no /dev entry anywhere the
adapter owns. The claim states which workload owns the radio, and it
holds that workload on the machine the radio is in. A stack that drives the
radio through the kernel opens its socket with the capabilities of its
own container, because a claim grants no privilege.
/dev/uhid belongs to the adapter’s claim because it is the kernel’s
inlet for a HID stack that runs in userspace. BlueZ sends HID over
GATT there and writes this node to present a BLE peripheral as an
input device, while classic Bluetooth HID passes through hidp inside the
kernel and needs no node. Declare the uhid module on a machine
whose adapter serves BLE input devices. Without it, such a device
pairs and connects, but no input node ever appears, and bluetoothd
logs input-hog profile accept failed.
liken stops the delivery walk at a bluetooth subtree. The kernel
puts the HID device of a connected peripheral under the adapter’s own
USB interface, so a game controller’s /dev/hidraw* node appears in
the adapter’s part of sysfs. That node belongs to the controller, and
a claim on the adapter does not deliver it. The same node moves when
BlueZ changes how it drives the kernel. With /dev/uhid present,
BlueZ 5.73 and later create the HID device under
/sys/devices/virtual/misc/uhid, where nothing connects it to the
adapter.
The input event nodes are delivered by number instead. The kernel
fixes the numbers of the first 32 event devices: character major 13,
minors 64 to 95, at /dev/input/event0 to /dev/input/event31. A
node in that range exists only while a device is registered at that
minor, and a peripheral that connects over the air registers one
after the claim was prepared. So the claim states each of the 32
nodes by its numbers, the container runtime creates them with mknod
and allows them in the device cgroup, and a program in the container
opens a controller’s node the moment the kernel registers it. The
Bluetooth stack uses those nodes together with /dev/uinput, the
kernel’s inlet for a virtual input device: it relays a controller’s
events from the real node, which comes and goes with the radio link,
into a virtual device whose node stays. uinput is built into the
liken kernel, so every machine has /dev/uinput. The bound is 32
input devices on one machine.
Serial-line adapters
Some USB devices present a serial line, and their kernel driver binds
only after a program attaches the line to the kernel’s serio layer
and keeps it attached. A USB-CEC adapter is one: cdc_acm creates
/dev/ttyACM0, and the adapter’s driver, pulse8_cec or
rainshadow_cec, binds to the serio port that exists only while the
attachment holds. The kernel’s
CEC admin guide
describes the setup on a general-purpose distribution. On liken, a
spec.serio
entry declares
the attachment, and the machine holds it for the life of the boot.
Load the drivers for a machine’s hardware
gives the steps.
The attached driver creates a CEC adapter, /dev/cec0, and the CEC
core registers an input device for the TV remote’s keys, with an
event node under /dev/input. liken publishes the adapter’s
serial-line interface as two devices:
- The CEC bus, with the suffix
-cecandsubsystem: cec. A claim on it delivers the/dev/cec*node. It is exclusive. The CEC core lets several processes open one adapter, but only one can be the exclusive initiator or follower, and two programs that configure logical addresses on one adapter undo each other. - The TV remote, with the suffix
-inputandsubsystem: input. A claim on it delivers the event node. It is exclusive, because one reader must own the remote’s key presses.
The two devices answer different claims: a program that talks on the
bus claims the first, and a program that reads the remote claims the
second. Both carry the adapter’s attributes, so one claim can pair
them with matchAttribute: liken.sh/address. The remote-control core
registers no LIRC node for a device whose only protocol is CEC, so a
claim delivers none.
The tty is never published, before or after the attachment. A pod
that received it could set another line discipline or write to the
line, and either one ends the attachment under every other claim.
So before the attachment, the adapter’s serial line publishes
nothing. Neither device keeps the interface’s bare name. Before an
entry declared the adapter, the bare name was the tty’s device, and
its claim delivered the tty and the usbfs node. After the entry, the
node rewrites such a claim’s specification to name a node that does
not exist, so its container cannot start again, and it never receives
the CEC node. Delete the claim, and claim the -cec or -input
device instead.
The usbfs node is not delivered either, because through it a program
could detach cdc_acm or reset the adapter and end the attachment
the same way. The adapter’s other interfaces keep their own devices
without the usbfs node: the Pulse-Eight’s HID interface still
publishes its hidraw and input nodes.
An unplug removes both devices from the slice. A pod that holds them
keeps file descriptors for nodes that are gone, and Kubernetes does
not evict it. A program that holds one of these claims must end when
a read or an ioctl returns ENODEV. The kubelet then restarts the
container, and the new container receives the nodes that the claim’s
CDI specification names at that moment.
While a device is absent, its claim’s specification names a node that does not exist, so a container cannot start. The old nodes would be unsafe: an event node carries a fixed device number, and another input device can take that number while the adapter is gone. The kubelet retries the container under its restart backoff, which doubles from 10 seconds to 5 minutes. The node rewrites every claim’s specification on each pass, so when the adapter returns to the same USB port, the devices return with the same names, and the container starts at its next retry with no new allocation. After a long absence, that retry can come up to 5 minutes after the adapter returns. A published name follows the USB port, so an adapter moved to another port is a different device, and its claimant needs a new pod.
Limits
- One slice for each node, with a maximum of 128 devices. If a node has more devices, it prints its device count on its console and drops the overflow, and no claim can reach the dropped devices. It does not divide the pool.
likensupplies no DeviceClasses. A DeviceClass states what a deployment needs, so each deployment writes its own. The guide has examples to start from.- Devices have no capacity and no taints at this time. If a device fails, it disappears from the slice at the next pass. This stops new allocations, but it does not change an allocation that is in use.
likendoes not keep a claim across a reboot. The claim is a Kubernetes object, and Kubernetes reschedules the pod that holds it together with the claim.
Hardware that is not published
You cannot claim hardware that has no driver bound to it, except a
whole USB device with no driver on any interface. liken shows this
hardware in two other places:
- The Machine’s status lists it as unclaimed hardware, with the modules that can drive it.
- The hardware report
lists it before the first install, as commented lines under
spec.modules.
To publish the device, declare the module in spec.modules. This is
the only necessary step. The device appears in the node’s slice at
the next reconcile pass.