Troubleshoot
By symptom. Each entry says how it shows, why, and what to do. The status fields and conditions referred to are listed in full in NRCell Status and Conditions; where the logs and traces of every component land is Read Logs and Traces.
A Cell Is in Phase Error
kubectl -n racora-system get nrcell
kubectl -n racora-system get nrcell <name> -o jsonpath='{.status.conditions}{"\n"}'
kubectl -n racora-system describe nrcell <name> # Events, for an identity mismatch
| Condition and reason | Cause | Fix |
|---|---|---|
ConfigGenerated=False, PlmnNotServed or TacNotServed | the cell's spec.plmn or spec.tac is outside the network identity (global.network); the core would not serve it and the CU-CP does not advertise it | change the cell, or change the identity in one place (Configure the 5G Core) |
PCIAllocated=False, TempPCIPoolExhausted | more than six cells without a spec.pci at once; the temporary pool is 1002 to 1007 | set spec.pci on some cells |
IdentityAllocated=False, SectorIdSpaceExhausted | every sector id under the gNB is taken | delete cells |
A cell can also carry ReportConfigConflict=True with reason DuplicateReportCfgId while its phase stays ConfigGenerated and its DU keeps running: two cells define the same mobility.reportConfigs id differently; the lexicographically first cell name's definition wins. The fix: make the definitions agree (Configure Mobility and Cell State).
A ruType other than uhd, zmq or dummy never gets this far: the API server rejects
the cell at apply time. Correcting the spec clears ConfigGenerated=False on the next
reconcile. PCIAllocated=False and IdentityAllocated=False are not cleared by the
controller, and the failed reconcile does not retry by itself: once you freed a PCI or an
identity, touch the cell (an annotation is enough) or restart the controller to reconcile
it again, and expect the old condition to stay on the object.
A Cell Keeps Its Temporary PCI
kubectl get nrcell shows a PCI between 1002 and 1007 with PCI-SOURCE controller-temp
and it never changes. The Intelligence Plane replaces a temporary PCI with a planned one,
so this means CU-IP is not running or not deciding. Check the cuip pod (next entry) and
the decision loop.
CU-IP Stays Pending
kubectl -n centralized-unit get pods
kubectl -n centralized-unit describe pod -l app=cuip | tail -n 12
The scheduler reports that no node matches the pod's node selector, or that no node has
nvidia.com/gpu to give. CU-IP requests one GPU and schedules only on a node labelled
racora.io/gpu-ready=true. Check the label and the allocatable GPU:
kubectl get nodes -L racora.io/gpu-ready
kubectl get node <node> -o jsonpath='{.status.allocatable.nvidia\.com/gpu}{"\n"}'
A missing label means the host's capabilities file did not report gpu_ready=true when
the node registered; fix the driver, then relabel the node
(Relabel a Node).
The DU Pod Is Pending
kubectl -n distributed-unit describe pod <du-pod> says no node satisfies the selector, or
Insufficient ettus.com/usrp. A real-radio DU needs a node with racora.io/rf-ready=true
(plus the hostname you named in duAssignment.nodeName) and one free USRP there; the
device plugin advertises only a radio that is present, and one DU holds one radio.
kubectl get nodes -L racora.io/rf-ready
kubectl get node <node> -o jsonpath='{.status.allocatable.ettus\.com/usrp}{"\n"}'
A node that shows 0 right after a kubelet restart is the next entry.
The USRP Is Not Advertised after a kubelet Restart
The node shows ettus.com/usrp: 0 after a kubelet or k3s-agent restart underneath the
device plugin: it registers with the kubelet once, at start-up. A running DU keeps its
radio; a new one does not schedule. Restart the plugin pod on that node, also after an
in-place upgrade that restarted the k3s agent on a worker:
kubectl -n racora-system delete pod -l app=ettus-device-plugin --field-selector spec.nodeName=<node>
The DU Restarts on Could not lock reference GPS time source
kubectl -n distributed-unit logs deploy/du-<cell> -c du --tail=20
The cell declares radioBackend.uhd.clockSource: gpsdo and the GPSDO has no fix, so the
DU exits and Kubernetes restarts it with growing backoff. A cold receiver takes minutes.
For a lone cell that does not need frame alignment, switch it to the internal clock:
kubectl -n racora-system patch nrcell <cell> --type merge -p '{"spec":{"radioBackend":{"uhd":{"clockSource":"internal"}}}}'
Checking lock and what sync needs is Connect a Radio.
The DU Logs F1-C: Failed to connect to CU-CP ... Connection refused
The DU is up but the CU-CP is not listening on F1-C (port 38472). Its log says why: a
CU-CP that cannot reach the AMF exits with CU-CP failed to connect to AMF and is
restarted until the core answers, and a CU-CP whose DUs all vanished at once (a radio
node lost) can stay up without accepting F1 again. Restart it; the DU reconnects by
itself:
kubectl -n centralized-unit get pods -l app=cu-cp
kubectl -n centralized-unit logs deploy/cu-cp -c cu-cp --tail=20
kubectl -n centralized-unit rollout restart deploy/cu-cp
The CU-CP Logs N2: Failed to connect to AMF
The CU-CP cannot reach the core over NGAP. With the Open5GS provider the AMF is the pod in
5g-core behind the amf Service; with an external core it is global.core.amfAddr.
kubectl -n 5g-core get pods,svc
kubectl -n 5g-core logs deploy/open5gs --tail=50
kubectl -n racora-system get deploy racora-controller -o jsonpath='{.spec.template.spec.containers[0].env[?(@.name=="RACORA_AMF_ADDR")].value}{"\n"}'
An external core must serve the network identity the CU-CP advertises (global.network);
NG Setup fails otherwise (External Core).
No Handover Happens
Two cells are on the air, the phone is attached to one, and walking toward the other changes nothing. Check, in order:
-
Frame alignment. Both cells run
clockSource: gpsdoand both radios hold GPS lock; neither DU is restarting. Free-running cells cannot see each other at all (How Mobility Reaches the CU-CP). -
Neighbors both ways. Each cell lists the other in its applied neighbors:
kubectl get nrcell -n racora-system -o custom-columns=NAME:.metadata.name,NEIGHBORS:.status.neighbors[*].nrCellRef -
The phone has the new configuration. A UE picks up a neighbor at its next RRC reconfiguration; a phone that attached before the relation was declared needs churn, such as toggling mobile data or re-attaching.
-
Measurement-triggered handover is on.
racora-controller.mobility.triggerHandoverFromMeasurements(defaulttrue) takes effect at the next CU-CP restart.
The CU-CP log shows how far it got:
kubectl exec -n centralized-unit deploy/cu-cp -c cu-cp -- grep -i handover /tmp/cu_cp.log | tail -20
The Phone Does Not Attach
Check the chain in order: the cell is on the air, the subscriber is provisioned, the core can read it.
-
kubectl -n racora-system get nrcellshowsConfigGeneratedand the DU pod isRunningand connected (the two entries above). -
kubectl -n racora-system get simshows the IMSIProvisioned(A subscriber is notProvisioned). -
The core's log for the IMSI. With Open5GS:
kubectl -n 5g-core logs deploy/open5gs --since=5m | grep <imsi>Registration completefollowed by a PDU session line is success.Cannot find SUPI in DBfrom the UDR means the core cannot read the subscriber; declare it as aSubscriber(Add Subscribers) so the adapter writes it in the shape the core reads.Registration rejectwith cause 7 ("5GS services not allowed", which Open5GS sends for a subscriber it cannot find) is what the phone got, and a phone caches that cause: after the fix, make it reselect the network (airplane mode is often not enough; a reboot or a manual network selection is).
The Phone Attaches but Has No Data
Registration and the PDU session succeed; nothing flows. With the Open5GS provider the UE
pool leaves the node through the egress-nat host unit: check that it is active, that
RACORA_UE_CIDR (/etc/racora/core-support.env) contains the chart's ueIpBase, and
that the CNI did not flush the masquerade rule (Open5GS, and step 8
of Install onto an Existing Kubernetes Cluster on a cluster you run). With an
external core the UPF must reach the CU-UP's pod IP over N3, GTP-U on UDP 2152, and the
CU-UP the UPF (External Core).
systemctl status egress-nat.service # on the control node: active (exited)?
cat /etc/racora/core-support.env # RACORA_UE_CIDR, default 10.45.0.0/16
kubectl -n 5g-core get cm open5gs-env -o jsonpath='{.data.UE_IP_BASE}{"\n"}' # must lie inside it
sudo iptables -t nat -S POSTROUTING | grep -F "$(sed -n 's/^RACORA_UE_CIDR=//p' /etc/racora/core-support.env)" # the masquerade rule
kubectl -n centralized-unit get pod -l app=cu-up -o wide # external core: the pod IP your UPF must reach
A Subscriber Is Not Provisioned
kubectl -n racora-system get sim
kubectl -n racora-system get sim <name> -o jsonpath='{.status.conditions[0].reason}: {.status.conditions[0].message}{"\n"}'
kubectl -n racora-system logs deploy/racora-core-controller --tail=50
The core controller's logs are only there: it sends no telemetry, so nothing of it is in ClickHouse.
| Phase and reason | Cause | Fix |
|---|---|---|
Pending, ProviderNotReady | the provider's declaration (ConfigMap racora-core-provider in 5g-core) is not there yet, or no running core pod matches its selector; retried every 30 s | wait for the core pod; kubectl -n 5g-core get pods,cm |
Pending, AdapterFailed | the provider's adapter could not run or exited non-zero; the message carries its output; retried every 30 s | read the message; the core pod's log has the rest |
Error, AdapterFailed | the provider's declaration is malformed | fix the ConfigMap (a provider module's bug) |
Error, SecretMissing | the Secret named by credentialsSecretRef is missing, or lacks k and opc/op, or a value is not 32 hex digits (amf: 4) | fix the Secret; the next reconcile provisions |
Unmanaged, Unmanaged | the selected provider has no subscriber adapter (an external core) | provision the SIM in your core |
Every subscriber is re-applied every 300 s (racora-controller.coreController.resyncSeconds),
so a fixed provider converges without you touching the Subscriber.
A k3s Upgrade Did Not Apply
Re-running the installer on the control node rewrites the HelmChart; k3s's helm-controller then
runs a job. If the release revision does not move:
kubectl -n kube-system get job -l helmcharts.helm.cattle.io/chart=racora
kubectl -n kube-system logs -l helmcharts.helm.cattle.io/chart=racora --tail=40
A render refusal (a retired values key, an identity the core cannot serve) is printed there
with its reason, and the previous release stays deployed and untouched. Fix the values in
the HelmChartConfig and the job re-runs (Upgrade).
Everything Is Running and a Cell Is Silent
Whether the cell is on the air is the CU-CP's answer, not Kubernetes': ask it with the
cell's decimal NCI over the Runtime Commands
(cell_status). operational_state: disabled with admin_state: locked is a cell you
declared adminState: Locked, or one the controller has not unlocked yet after a PCI
retune (Configure Mobility and Cell State).