The Node Contract
Racora's network layer (the chart and the controller) is identical on every
Kubernetes. What it needs from the nodes underneath is this contract. The
k3s platform implements all of it during install; on a cluster you run
(INSTALL_RACORA_PLATFORM=kubernetes) you implement it with your own tooling,
using Install onto an Existing Kubernetes Cluster. Everything below is
what the code keys on.
Labels
Written by the provisioner's validator into /etc/racora/node-capabilities (with
kernel, isolated_cpus, fronthaul_nic, fronthaul_phc and evaluated_at) and
applied to the Node as labels (Node Provisioner).
| Label | true means the node passed the gate for… | Consumed by |
|---|---|---|
racora.io/rf-ready | RT kernel and isolated CPUs and the racora TuneD profile and UHD installed | the controller: every DU with radioBackend.ruType: uhd gets nodeSelector racora.io/rf-ready: "true"; virtual cells (zmq, dummy) get no such selector |
racora.io/gpu-ready | the NVIDIA driver works on the host (nvidia-smi) and the container toolkit is installed; the container-level proof is the NVIDIA device plugin once the node is in the cluster | the chart: CU-IP merges racora.io/gpu-ready: "true" into its nodeSelector when cuIp.gpuCount > 0 |
racora.io/fronthaul-ready | a NIC with hardware timestamping + PHC and PTP running and the RT-latency stage holding (no deep idle states on the isolated cores, the NIC's IRQs off them) | no shipped workload selects it; the platforms apply it from the capabilities file like the other two |
k3s records a node's labels at the agent's first registration, so on the k3s platform a host is provisioned before it joins and its labels are written into a k3s config drop-in.
A label states that the host passed the gate at evaluated_at; the cluster does
not re-evaluate it. Re-run validate-rf-node.sh (and re-label) after changing
the host. spec.duAssignment.nodeName on an NRCell is merged as a
kubernetes.io/hostname selector next to rf-ready, never a raw
pod.spec.nodeName bind, so the scheduler still checks the node.
Device-Plugin Resources
| Resource | Advertised by | Requested by |
|---|---|---|
ettus.com/usrp | DaemonSet ettus-device-plugin (racora-node chart, racora-system) | every UHD DU, limits: ettus.com/usrp: 1 |
nvidia.com/gpu | DaemonSet nvidia-device-plugin-daemonset (racora-node) — runs under RuntimeClass/nvidia | CU-IP, cuIp.gpuCount |
The plugins ship with the chart on every platform. RuntimeClass/nvidia is
provided by the platform (k3s registers it when the NVIDIA container toolkit is
present) or created by the chart (racora-node.nvidia.runtimeClass.create=true)
— the containerd nvidia handler on GPU nodes is a host prerequisite either way.
What the Controller Bakes into Every DU
- Capabilities
SYS_NICE,IPC_LOCK,PERFMON(the gNB binaries carry file capabilities; real-time threads needSYS_NICE). - UHD DUs read the live isolated-CPU set from
/sys/devices/system/cpu/isolatedat pod start andtasksetthe DU onto it; fallbackISOLATED_CPUSon the controller; with neither, the DU runs unpinned. Isolation is therefore a per-node property, discovered at runtime. - UHD DUs use
strategy: Recreate(the USRP is exclusive; a surge pod cannot obtain it while the old pod holds it). - The B210 must be on USB 3 (bandwidth); a stale USB claim after a crash is cleared by deleting the DU pod.
imagePullPolicy: IfNotPresenteverywhere: images come from the racora registry namespace (re-rootable viaglobal.systemDefaultRegistry) or from a preloaded air-gap archive.
What the Cluster Must Provide
The cluster-side requirements are on Requirements; the by-hand steps for a cluster you run are Install onto an Existing Kubernetes Cluster. In one line each:
- the
sctpkernel module on every node (F1-C, E1 and NGAP are SCTP over ClusterIP Services); - one schedulable node with the GPU for the CU planes, an in-cluster core and the
racora-systemcontrollers (controlPlanePin, or yournodeSelectorwhere the control plane is tainted); - a default StorageClass: the Open5GS provider's subscriber claim and the optional
monitoring claims set no
storageClassName, so the default class provisions them; - the privileged namespaces allowed (CU-CP, CU-UP and every DU add capabilities; the
Open5GS provider runs
hostNetworkandprivileged; the USRP plugin is privileged); - the selected core provider's host units on the control node (Open5GS).
Who Implements What
Each platform binds the contract its own way; the side-by-side is the Platforms section.