Skip to main content

The Node Contract

Racora's network layer (the chart and the controller) is identical on every Kubernetes. What it needs from the nodes underneath is this contract. The k3s platform implements all of it during install; on a cluster you run (INSTALL_RACORA_PLATFORM=kubernetes) you implement it with your own tooling, using Install onto an Existing Kubernetes Cluster. Everything below is what the code keys on.

Labels​

Written by the provisioner's validator into /etc/racora/node-capabilities (with kernel, isolated_cpus, fronthaul_nic, fronthaul_phc and evaluated_at) and applied to the Node as labels (Node Provisioner).

Labeltrue means the node passed the gate for…Consumed by
racora.io/rf-readyRT kernel and isolated CPUs and the racora TuneD profile and UHD installedthe controller: every DU with radioBackend.ruType: uhd gets nodeSelector racora.io/rf-ready: "true"; virtual cells (zmq, dummy) get no such selector
racora.io/gpu-readythe NVIDIA driver works on the host (nvidia-smi) and the container toolkit is installed; the container-level proof is the NVIDIA device plugin once the node is in the clusterthe chart: CU-IP merges racora.io/gpu-ready: "true" into its nodeSelector when cuIp.gpuCount > 0
racora.io/fronthaul-readya NIC with hardware timestamping + PHC and PTP running and the RT-latency stage holding (no deep idle states on the isolated cores, the NIC's IRQs off them)no shipped workload selects it; the platforms apply it from the capabilities file like the other two

k3s records a node's labels at the agent's first registration, so on the k3s platform a host is provisioned before it joins and its labels are written into a k3s config drop-in.

A label states that the host passed the gate at evaluated_at; the cluster does not re-evaluate it. Re-run validate-rf-node.sh (and re-label) after changing the host. spec.duAssignment.nodeName on an NRCell is merged as a kubernetes.io/hostname selector next to rf-ready, never a raw pod.spec.nodeName bind, so the scheduler still checks the node.

Device-Plugin Resources​

ResourceAdvertised byRequested by
ettus.com/usrpDaemonSet ettus-device-plugin (racora-node chart, racora-system)every UHD DU, limits: ettus.com/usrp: 1
nvidia.com/gpuDaemonSet nvidia-device-plugin-daemonset (racora-node) — runs under RuntimeClass/nvidiaCU-IP, cuIp.gpuCount

The plugins ship with the chart on every platform. RuntimeClass/nvidia is provided by the platform (k3s registers it when the NVIDIA container toolkit is present) or created by the chart (racora-node.nvidia.runtimeClass.create=true) — the containerd nvidia handler on GPU nodes is a host prerequisite either way.

What the Controller Bakes into Every DU​

  • Capabilities SYS_NICE, IPC_LOCK, PERFMON (the gNB binaries carry file capabilities; real-time threads need SYS_NICE).
  • UHD DUs read the live isolated-CPU set from /sys/devices/system/cpu/isolated at pod start and taskset the DU onto it; fallback ISOLATED_CPUS on the controller; with neither, the DU runs unpinned. Isolation is therefore a per-node property, discovered at runtime.
  • UHD DUs use strategy: Recreate (the USRP is exclusive; a surge pod cannot obtain it while the old pod holds it).
  • The B210 must be on USB 3 (bandwidth); a stale USB claim after a crash is cleared by deleting the DU pod.
  • imagePullPolicy: IfNotPresent everywhere: images come from the racora registry namespace (re-rootable via global.systemDefaultRegistry) or from a preloaded air-gap archive.

What the Cluster Must Provide​

The cluster-side requirements are on Requirements; the by-hand steps for a cluster you run are Install onto an Existing Kubernetes Cluster. In one line each:

  • the sctp kernel module on every node (F1-C, E1 and NGAP are SCTP over ClusterIP Services);
  • one schedulable node with the GPU for the CU planes, an in-cluster core and the racora-system controllers (controlPlanePin, or your nodeSelector where the control plane is tainted);
  • a default StorageClass: the Open5GS provider's subscriber claim and the optional monitoring claims set no storageClassName, so the default class provisions them;
  • the privileged namespaces allowed (CU-CP, CU-UP and every DU add capabilities; the Open5GS provider runs hostNetwork and privileged; the USRP plugin is privileged);
  • the selected core provider's host units on the control node (Open5GS).

Who Implements What​

Each platform binds the contract its own way; the side-by-side is the Platforms section.