Claim a GPU by what it decodes
A liken render node states the GPU’s identity, and nothing about
what its media driver can do. On a fleet of mixed GPUs, a claim on any
render node can land on a GPU whose driver cannot decode the file’s
codec, or cannot scale its 10-bit frames. The workload then falls
back to a slower path, or fails.
The capabilities agent of media-operator publishes what each GPU’s
driver states, as a media.liken.sh device for each render node. A
claim asks for that device beside the render node, and a constraint
keeps the two on one GPU. The scheduler then places the pod only on a
GPU whose driver states what the claim asks for.
You need:
media-operatorinstalled with the capabilities agent (Install the operator ).- A
likenrelease that publishesresource.kubernetes.io/pciBusIDon its render nodes. On an older release, the claim below never allocates.
1. See what each GPU states
kubectl get resourceslices
kubectl get resourceslice <node>-media.liken.sh -o yaml
Each device is one GPU:
- name: pci-0000-00-02-0
allowMultipleAllocations: true
attributes:
address: {string: "0000:00:02.0"}
decodeAV1Main: {bool: true}
decodeAV1Main10: {bool: true}
decodeH264: {bool: true}
decodeHEVCMain: {bool: true}
decodeHEVCMain10: {bool: true}
decodeMPEG2: {bool: true}
decodeVC1: {bool: false}
decodeVP8: {bool: true}
decodeVP9: {bool: true}
decodeVP9Profile2: {bool: true}
driver: {string: i915}
encodeH264: {bool: true}
encodeH264LowPower: {bool: false}
encodeHEVCMain: {bool: true}
encodeHEVCMain10: {bool: true}
encodeHEVCMain10LowPower: {bool: false}
encodeHEVCMainLowPower: {bool: false}
name: {string: Meteor Lake-P [Intel Arc Graphics]}
product: {string: 7d55}
resource.kubernetes.io/pciBusID: {string: "0000:00:02.0"}
scale10bit: {bool: true}
scale8bit: {bool: true}
vaDriver: {string: Intel iHD driver for Intel(R) Gen Graphics - 26.1.2 ()}
vendor: {string: "8086"}
Render node capabilities describes each attribute, and where its value comes from.
2. Write the classes
media-operator ships no class for a capability. A class encodes a
deployment’s purposes, so you write the classes that your workloads
claim through. Two classes are enough here.
The first selects the media.liken.sh devices of GPUs that decode
10-bit HEVC and scale 10-bit frames on the video processor, which is
the whole path of a 10-bit film from the decoder to a scaled frame:
apiVersion: resource.k8s.io/v1
kind: DeviceClass
metadata:
name: decode-10bit
spec:
selectors:
- cel:
expression: |
device.driver == "media.liken.sh" &&
device.attributes["media.liken.sh"].decodeHEVCMain10 &&
device.attributes["media.liken.sh"].scale10bit
The second selects the render nodes that liken publishes. If you
have a class like it, such as display-render from
display-operator, use that class and skip this one. It works the
same way:
apiVersion: resource.k8s.io/v1
kind: DeviceClass
metadata:
name: gpu-render
spec:
selectors:
- cel:
expression: |
device.driver == "liken.sh" &&
has(device.attributes["liken.sh"].renderNode)
media-render is not for workloads. It is the class of the agent’s
own claim, and the install ships it as wiring.
3. Claim the render node and the capability together
The claim has two requests: the render node from liken, and the
media.liken.sh device of a GPU that states what the workload needs.
The constraint requires both devices to have the same
resource.kubernetes.io/pciBusID, which is the GPU’s PCI address, so
both are the same GPU:
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
name: decode-10bit
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu-render
- name: decodes
exactly:
deviceClassName: decode-10bit
constraints:
- requests: [gpu, decodes]
matchAttribute: resource.kubernetes.io/pciBusID
The pod names the claim, and the container receives the render node of the GPU the scheduler chose:
spec:
resourceClaims:
- name: gpu
resourceClaimTemplateName: decode-10bit
containers:
- name: worker
resources:
claims:
- name: gpu
The media.liken.sh device delivers nothing into the container. It
is a statement about the GPU, and any number of claims can allocate
it at once. The render node is shareable too, so several workloads
decode on one GPU at the same time.
4. Ask for another set
For a set that decode-10bit does not state, add a selector to the
request. It narrows the class, so it works on any class that selects
media.liken.sh devices. This request asks for a GPU that decodes VP9
and scales 8-bit frames, through a class of your own that selects
every media.liken.sh device (device.driver == "media.liken.sh"),
here named media-gpu:
- name: decodes
exactly:
deviceClassName: media-gpu
selectors:
- cel:
expression: |
device.attributes["media.liken.sh"].decodeVP9 &&
device.attributes["media.liken.sh"].scale8bit
A selector on decode-10bit works the same way and requires the
10-bit capabilities as well. To reuse a set, write another class with
the expression in it.
Every device has every capability attribute, true or false, so a
selector needs no has() check.
5. Fall back when no GPU qualifies
A claim that no GPU satisfies stays unallocated, and its pod stays
Pending:
kubectl describe pod <pod>
The events name the claim that cannot allocate. Such a workload has two ways forward: a second claim on any render node, which the workload uses with its own fallback path, or no claim at all and work on the CPU. The scheduler does not choose between them; the workload’s controller does.
When a GPU states false for a capability it has
The values are the driver’s own statements, read through libva. The
agent does not test them. A capability that the GPU has but the
driver in the agent’s image does not list publishes as false. The
agent’s log has the driver’s name and every value for each GPU:
kubectl -n liken-system logs ds/media-capabilities
A query that failed has the failure there, word for word, and the
device then has every capability false and no vaDriver.