Skip to main content

Troubleshoot

By symptom. Each entry says how it shows, why, and what to do. The status fields and conditions referred to are listed in full in NRCell Status and Conditions; where the logs and traces of every component land is Read Logs and Traces.

A Cell Is in Phase Error​

kubectl -n racora-system get nrcell
kubectl -n racora-system get nrcell <name> -o jsonpath='{.status.conditions}{"\n"}'
kubectl -n racora-system describe nrcell <name> # Events, for an identity mismatch
Condition and reasonCauseFix
ConfigGenerated=False, PlmnNotServed or TacNotServedthe cell's spec.plmn or spec.tac is outside the network identity (global.network); the core would not serve it and the CU-CP does not advertise itchange the cell, or change the identity in one place (Configure the 5G Core)
PCIAllocated=False, TempPCIPoolExhaustedmore than six cells without a spec.pci at once; the temporary pool is 1002 to 1007set spec.pci on some cells
IdentityAllocated=False, SectorIdSpaceExhaustedevery sector id under the gNB is takendelete cells

A cell can also carry ReportConfigConflict=True with reason DuplicateReportCfgId while its phase stays ConfigGenerated and its DU keeps running: two cells define the same mobility.reportConfigs id differently; the lexicographically first cell name's definition wins. The fix: make the definitions agree (Configure Mobility and Cell State).

A ruType other than uhd, zmq or dummy never gets this far: the API server rejects the cell at apply time. Correcting the spec clears ConfigGenerated=False on the next reconcile. PCIAllocated=False and IdentityAllocated=False are not cleared by the controller, and the failed reconcile does not retry by itself: once you freed a PCI or an identity, touch the cell (an annotation is enough) or restart the controller to reconcile it again, and expect the old condition to stay on the object.

A Cell Keeps Its Temporary PCI​

kubectl get nrcell shows a PCI between 1002 and 1007 with PCI-SOURCE controller-temp and it never changes. The Intelligence Plane replaces a temporary PCI with a planned one, so this means CU-IP is not running or not deciding. Check the cuip pod (next entry) and the decision loop.

CU-IP Stays Pending​

kubectl -n centralized-unit get pods
kubectl -n centralized-unit describe pod -l app=cuip | tail -n 12

The scheduler reports that no node matches the pod's node selector, or that no node has nvidia.com/gpu to give. CU-IP requests one GPU and schedules only on a node labelled racora.io/gpu-ready=true. Check the label and the allocatable GPU:

kubectl get nodes -L racora.io/gpu-ready
kubectl get node <node> -o jsonpath='{.status.allocatable.nvidia\.com/gpu}{"\n"}'

A missing label means the host's capabilities file did not report gpu_ready=true when the node registered; fix the driver, then relabel the node (Relabel a Node).

The DU Pod Is Pending​

kubectl -n distributed-unit describe pod <du-pod> says no node satisfies the selector, or Insufficient ettus.com/usrp. A real-radio DU needs a node with racora.io/rf-ready=true (plus the hostname you named in duAssignment.nodeName) and one free USRP there; the device plugin advertises only a radio that is present, and one DU holds one radio.

kubectl get nodes -L racora.io/rf-ready
kubectl get node <node> -o jsonpath='{.status.allocatable.ettus\.com/usrp}{"\n"}'

A node that shows 0 right after a kubelet restart is the next entry.

The USRP Is Not Advertised after a kubelet Restart​

The node shows ettus.com/usrp: 0 after a kubelet or k3s-agent restart underneath the device plugin: it registers with the kubelet once, at start-up. A running DU keeps its radio; a new one does not schedule. Restart the plugin pod on that node, also after an in-place upgrade that restarted the k3s agent on a worker:

kubectl -n racora-system delete pod -l app=ettus-device-plugin --field-selector spec.nodeName=<node>

The DU Restarts on Could not lock reference GPS time source​

kubectl -n distributed-unit logs deploy/du-<cell> -c du --tail=20

The cell declares radioBackend.uhd.clockSource: gpsdo and the GPSDO has no fix, so the DU exits and Kubernetes restarts it with growing backoff. A cold receiver takes minutes. For a lone cell that does not need frame alignment, switch it to the internal clock:

kubectl -n racora-system patch nrcell <cell> --type merge -p '{"spec":{"radioBackend":{"uhd":{"clockSource":"internal"}}}}'

Checking lock and what sync needs is Connect a Radio.

The DU Logs F1-C: Failed to connect to CU-CP ... Connection refused​

The DU is up but the CU-CP is not listening on F1-C (port 38472). Its log says why: a CU-CP that cannot reach the AMF exits with CU-CP failed to connect to AMF and is restarted until the core answers, and a CU-CP whose DUs all vanished at once (a radio node lost) can stay up without accepting F1 again. Restart it; the DU reconnects by itself:

kubectl -n centralized-unit get pods -l app=cu-cp
kubectl -n centralized-unit logs deploy/cu-cp -c cu-cp --tail=20
kubectl -n centralized-unit rollout restart deploy/cu-cp

The CU-CP Logs N2: Failed to connect to AMF​

The CU-CP cannot reach the core over NGAP. With the Open5GS provider the AMF is the pod in 5g-core behind the amf Service; with an external core it is global.core.amfAddr.

kubectl -n 5g-core get pods,svc
kubectl -n 5g-core logs deploy/open5gs --tail=50
kubectl -n racora-system get deploy racora-controller -o jsonpath='{.spec.template.spec.containers[0].env[?(@.name=="RACORA_AMF_ADDR")].value}{"\n"}'

An external core must serve the network identity the CU-CP advertises (global.network); NG Setup fails otherwise (External Core).

No Handover Happens​

Two cells are on the air, the phone is attached to one, and walking toward the other changes nothing. Check, in order:

  1. Frame alignment. Both cells run clockSource: gpsdo and both radios hold GPS lock; neither DU is restarting. Free-running cells cannot see each other at all (How Mobility Reaches the CU-CP).

  2. Neighbors both ways. Each cell lists the other in its applied neighbors:

    kubectl get nrcell -n racora-system -o custom-columns=NAME:.metadata.name,NEIGHBORS:.status.neighbors[*].nrCellRef
  3. The phone has the new configuration. A UE picks up a neighbor at its next RRC reconfiguration; a phone that attached before the relation was declared needs churn, such as toggling mobile data or re-attaching.

  4. Measurement-triggered handover is on. racora-controller.mobility.triggerHandoverFromMeasurements (default true) takes effect at the next CU-CP restart.

The CU-CP log shows how far it got:

kubectl exec -n centralized-unit deploy/cu-cp -c cu-cp -- grep -i handover /tmp/cu_cp.log | tail -20

The Phone Does Not Attach​

Check the chain in order: the cell is on the air, the subscriber is provisioned, the core can read it.

  1. kubectl -n racora-system get nrcell shows ConfigGenerated and the DU pod is Running and connected (the two entries above).

  2. kubectl -n racora-system get sim shows the IMSI Provisioned (A subscriber is not Provisioned).

  3. The core's log for the IMSI. With Open5GS:

    kubectl -n 5g-core logs deploy/open5gs --since=5m | grep <imsi>

    Registration complete followed by a PDU session line is success. Cannot find SUPI in DB from the UDR means the core cannot read the subscriber; declare it as a Subscriber (Add Subscribers) so the adapter writes it in the shape the core reads. Registration reject with cause 7 ("5GS services not allowed", which Open5GS sends for a subscriber it cannot find) is what the phone got, and a phone caches that cause: after the fix, make it reselect the network (airplane mode is often not enough; a reboot or a manual network selection is).

The Phone Attaches but Has No Data​

Registration and the PDU session succeed; nothing flows. With the Open5GS provider the UE pool leaves the node through the egress-nat host unit: check that it is active, that RACORA_UE_CIDR (/etc/racora/core-support.env) contains the chart's ueIpBase, and that the CNI did not flush the masquerade rule (Open5GS, and step 8 of Install onto an Existing Kubernetes Cluster on a cluster you run). With an external core the UPF must reach the CU-UP's pod IP over N3, GTP-U on UDP 2152, and the CU-UP the UPF (External Core).

systemctl status egress-nat.service # on the control node: active (exited)?
cat /etc/racora/core-support.env # RACORA_UE_CIDR, default 10.45.0.0/16
kubectl -n 5g-core get cm open5gs-env -o jsonpath='{.data.UE_IP_BASE}{"\n"}' # must lie inside it
sudo iptables -t nat -S POSTROUTING | grep -F "$(sed -n 's/^RACORA_UE_CIDR=//p' /etc/racora/core-support.env)" # the masquerade rule
kubectl -n centralized-unit get pod -l app=cu-up -o wide # external core: the pod IP your UPF must reach

A Subscriber Is Not Provisioned​

kubectl -n racora-system get sim
kubectl -n racora-system get sim <name> -o jsonpath='{.status.conditions[0].reason}: {.status.conditions[0].message}{"\n"}'
kubectl -n racora-system logs deploy/racora-core-controller --tail=50

The core controller's logs are only there: it sends no telemetry, so nothing of it is in ClickHouse.

Phase and reasonCauseFix
Pending, ProviderNotReadythe provider's declaration (ConfigMap racora-core-provider in 5g-core) is not there yet, or no running core pod matches its selector; retried every 30 swait for the core pod; kubectl -n 5g-core get pods,cm
Pending, AdapterFailedthe provider's adapter could not run or exited non-zero; the message carries its output; retried every 30 sread the message; the core pod's log has the rest
Error, AdapterFailedthe provider's declaration is malformedfix the ConfigMap (a provider module's bug)
Error, SecretMissingthe Secret named by credentialsSecretRef is missing, or lacks k and opc/op, or a value is not 32 hex digits (amf: 4)fix the Secret; the next reconcile provisions
Unmanaged, Unmanagedthe selected provider has no subscriber adapter (an external core)provision the SIM in your core

Every subscriber is re-applied every 300 s (racora-controller.coreController.resyncSeconds), so a fixed provider converges without you touching the Subscriber.

A k3s Upgrade Did Not Apply​

Re-running the installer on the control node rewrites the HelmChart; k3s's helm-controller then runs a job. If the release revision does not move:

kubectl -n kube-system get job -l helmcharts.helm.cattle.io/chart=racora
kubectl -n kube-system logs -l helmcharts.helm.cattle.io/chart=racora --tail=40

A render refusal (a retired values key, an identity the core cannot serve) is printed there with its reason, and the previous release stays deployed and untouched. Fix the values in the HelmChartConfig and the job re-runs (Upgrade).

Everything Is Running and a Cell Is Silent​

Whether the cell is on the air is the CU-CP's answer, not Kubernetes': ask it with the cell's decimal NCI over the Runtime Commands (cell_status). operational_state: disabled with admin_state: locked is a cell you declared adminState: Locked, or one the controller has not unlocked yet after a PCI retune (Configure Mobility and Cell State).