Volumes

A person or an operator writes two objects. The PersistentVolume names the driver, the handle, the class, an access mode, a capacity, Retain, and a claimRef. The capacity is a hint, not a limit. The driver enforces no size limit. The claimRef binds the volume to one claim and no other. The claim names the class, the same access mode, and the same size. In this example the access mode is ReadWriteMany: many nodes mount this one volume read-write, and each node keeps its own copy.

apiVersion: v1
kind: PersistentVolume
metadata:
  name: example-store
spec:
  storageClassName: per-node
  accessModes: [ReadWriteMany]
  capacity:
    storage: 20Gi
  persistentVolumeReclaimPolicy: Retain
  claimRef:
    namespace: example
    name: store
  csi:
    driver: per-node.liken.sh
    volumeHandle: example-store
---
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: store
  namespace: example
spec:
  storageClassName: per-node
  accessModes: [ReadWriteMany]
  resources:
    requests:
      storage: 20Gi

The handle

The handle is spec.csi.volumeHandle. By convention it is the PersistentVolume’s name, so the name a person reads in kubectl get pv is the name of the directory on every node. The handle names the volume’s directory under the store on every node it has visited, so it is a DNS-1123 subdomain name: lower-case letters, digits, - and ., starting and ending with a letter or a digit, at most 253 characters, no /, and no empty segment between two dots. The driver refuses any other handle with InvalidArgument and posts a PerNodeVolumeRefused event on the pod.

The access mode

Use ReadWriteMany for a volume that pods on many nodes mount. Many nodes do mount one volume read-write, and each node holds its own copy. The driver enforces one pod per node per volume itself, and refuses a second pod on the same node with a PerNodeVolumeHeld event.

Use ReadWriteOncePod for a volume that only one pod in the whole cluster may mount, such as a database with one writer. Set it on the PersistentVolume and on the claim. The scheduler admits one pod of a ReadWriteOncePod claim in the cluster, and keeps every other pod of the claim Pending until that pod is gone. A pod that starts on a different node gets that node’s copy, not the copy on the node the last pod left. A Deployment with the RollingUpdate strategy cannot start a new pod while its old pod holds the claim, so give it the Recreate strategy.

Do not use ReadWriteOnce. It says that one node holds the volume, which is false, because every node keeps its own copy.

The node plugin declares the CSI capability SINGLE_NODE_MULTI_WRITER, so the kubelet sends the CSI access mode SINGLE_NODE_SINGLE_WRITER for a ReadWriteOncePod claim. The driver publishes every access mode the same way: it binds this node’s copy onto the pod’s target, read-only when the pod mounts the volume read-only.

What deleting removes

Delete the PersistentVolume to remove every copy. The node plugin on each node watches the PersistentVolumes of this driver, and it removes a copy whose handle no PersistentVolume names. It does so after its watch has synced, only while no pod on that node holds the copy, on the deletion itself and again on every --sweep-every tick. A copy a pod still holds stays until that pod is gone.

Deleting the claim alone leaves the volume Released and every copy in place, because no controller runs a deleter. Deleting a pod removes nothing: the copy outlives its pod.

Events

The driver posts a Warning event on the pod for every refusal, so kubectl describe pod says why a mount did not happen.

The kubelet retries a refused mount with a backoff that grows to about two minutes. Each retry that the driver refuses for the same reason, with the same message, within 10 minutes of the last one, increases the count of the event already posted and posts no new one. So a pod that waits an hour for a handle that another pod holds carries one line, such as (x37 over 1h). The API server deletes an event one hour after its last change.

Reason When
PerNodeVolumeRefused The handle is not a name the driver can put under the store. The message says which rule it broke.
PerNodeVolumeHeld Another pod on this node holds the handle. The message names that pod.
PerNodeMountFailed The driver could not make the copy, could not make the target, or could not bind one onto the other. The message carries the error.