Kubernetes Infrastructure
The device plugin runs as a DaemonSet on every RISC-V node. At startup it labels the node with the detected SoC board name, then registers with the kubelet to advertise a scheduling resource. Together, the label and the resource enable board-specific scheduling with exclusive node access.
Source: the runner/device-plugin/ directory
How it fits together
flowchart TD
subgraph Node["RISC-V Node"]
HWP["riscv_hwprobe syscall"]
DT["/sys/firmware/devicetree/base/compatible<br/>(Scaleway fallback)"]
DP["Device Plugin Pod"]
KL["Kubelet"]
RP["Runner Pod"]
end
subgraph Cluster["Kubernetes Cluster"]
SCHED["Scheduler"]
end
HWP <-->|"mvendorid, marchid, mimpid"| DP
DT <-->|"probe failed: read compatible"| DP
DP -->|"patch node label riseproject.dev/board=spacemit-k3"| SCHED
DP -->|"advertise riseproject.com/runner=1"| KL
KL -->|report allocatable resources| SCHED
KL -->|schedule| RP
SCHED -->|"nodeSelector: riseproject.dev/board\nlimits: riseproject.com/runner=1"| KL
Device plugin
The device plugin implements the Kubernetes Device Plugin API via gRPC. It advertises exactly one riseproject.com/runner resource per node.
Runner pods request this resource:
resources:
limits:
riseproject.com/runner: "1"
Since only one unit exists per node, the Kubernetes scheduler will never place two runner pods on the same node. This enforces exclusive access without taints or manual coordination.
How it works
- The plugin starts and registers with kubelet via a Unix socket at
/var/lib/kubelet/device-plugins/rise-riscv-runner.sock ListAndWatch()advertises a single healthy device (runner-0) to the kubeletAllocate()returns an empty response. No actual device allocation is needed; only the scheduling constraint matters- A file watcher monitors the kubelet socket directory and re-registers if kubelet restarts
Node labelling
The device plugin detects the SoC on each RISC-V node at startup and applies a riseproject.dev/board label. Runner pods use this label in their nodeSelector to land on the correct hardware.
The device plugin owns the board key only. Placement also uses riseproject.dev/provider, which records the vendor supplying the machine and is applied by hand when a node joins (see Cluster Provisioning). The plugin patches just the board key, so a hand-applied provider label is preserved across restarts. Nodes are selected on both labels together, so a node missing provider accepts no jobs.
SoC detection
The primary key is the riscv_hwprobe(2) syscall, which returns the hardware identity triple (mvendorid, marchid, mimpid) read from the CPU CSRs.
- Call
riscv_hwprobefor the three ID keys and log the triple as hex. - Match the triple against a hand-maintained list of known SoCs:
mvendorid | marchid | mimpid | Board label |
|---|---|---|---|
0x710 | 0x8000000058000001 | 0x1000000049772200 | spacemit-k1 |
0x710 | 0x8000000058000002 | 0x33d8a600 | spacemit-k3 |
0x710 | 0x8000000058000002 | 0x4c4d900 | spacemit-v100 |
0x5b7 | 0x8000000009140d00 | 0x100d000 | zhihe-a210 |
- If the triple matches no entry,
Detectreturns an error and the plugin exits. An unrecognized node fails loudly rather than mislabelling itself. To add the board, read the logged triple and append an entry to the list.
The Scaleway EM-RV1 and Zhihe A210 support device tree fallback: if riscv_hwprobe fails (or for vendor kernels without the syscall), detection falls back to reading /sys/firmware/devicetree/base/compatible; a scaleway,em-rv1 prefix yields scaleway-em-rv1, and zhihe,a210 yields zhihe-a210. On heterogeneous multi-cluster RISC-V SoCs like the Zhihe A210 (ESWIN EIC7700X, 4x P550 cores 0-3 + 4x custom cores 4-7), an all-CPU riscv_hwprobe(2) query returns -1 for keys that differ across clusters; probeHWID automatically queries CPU 0 to obtain the primary cluster triple. Any other board on a kernel without the syscall is treated as the original probe failure.
DaemonSet configuration
- Namespace:
kube-system - Node selector:
kubernetes.io/arch: riscv64 - Priority:
system-node-critical - RBAC: ServiceAccount with ClusterRole granting
getandpatchon nodes - Environment:
NODE_NAMEfrom downward API (spec.nodeName) - Volume mounts:
/var/lib/kubelet/device-plugins(host path),/sys(read-only host path) - Privileged: Yes (device tree access for the Scaleway fallback)
- Image:
ghcr.io/riseproject-dev/riscv-runner/device-plugin:prod
Source files
| File | Role |
|---|---|
cmd/k8s-device-plugin/main.go | Entry point: label node, then start device plugin |
pkg/plugin/plugin.go | gRPC server, kubelet registration, watchdog |
pkg/soc/detect.go | riscv_hwprobe triple matching, Scaleway device tree fallback, SoC → board mapping |
pkg/labeler/labeler.go | Kubernetes API node label patching |
k8s-ds-device-plugin.yaml | DaemonSet + RBAC manifest |