Upgrade Racora
A Racora release is one chart with one pinned image set. Upgrading means applying the
next release to the cluster you have: on the k3s platform by re-running the installer
on the control node, on a cluster you run with helm upgrade. Kubernetes then rolls exactly
the pods whose spec changed, on the nodes they run on. There is no RAN version per node:
the RAN is the control node's release, and a worker only ever runs the pods scheduled onto it.
k3s Platform
On the server only:
curl -sfL https://get.racora.io | INSTALL_RACORA_VERSION=vX.Y.Z sh -
The re-run keeps the core provider and AMF address recorded at install time
(/etc/racora/install.env) unless you give others; any INSTALL_K3S_* or K3S_* flag
you set at install must be given again, because k3s rewrites its service file from them.
The installer downloads the release's k3s binary only if the pinned upstream version
changed, re-renders the core provider's host units, restarts k3s when its binary or
service file changed, and only then stages
the payload: the CRDs are re-applied out-of-band (racora-crds.yaml) and the HelmChart
is rewritten with the new chart package. Same name, so k3s's helm-controller runs
helm upgrade, and the Deployments roll with the new images and configuration; a DU on
a worker pulls its new image from the registry.
Workers usually need nothing. Re-run node mode on a worker (with its K3S_URL,
K3S_TOKEN and the same INSTALL_RACORA_VERSION as the control node) only when the k3s
pin moved, which the k3s column of Release History shows: on an
already provisioned host it validates and
joins straight away, swapping the agent binary. After that restart, the USRP device
plugin on the node must re-register (Add a Radio Node).
A Cluster You Run
A plain helm upgrade with the two overlays from the release's installer package
(Install onto an Existing Kubernetes Cluster); no host is involved:
helm upgrade --install racora oci://registry.gitlab.com/cognitive-network-solutions/racora/charts/racora \
--namespace default --version X.Y.Z --reset-values \
-f installer/platforms/kubernetes/values.yaml \
-f installer/cores/<provider>/values.yaml \
-f my-values.yaml \
--set global.systemDefaultRegistry=<registry>
This is the platform's command with
--reset-values and the new release's installer/ overlays. Pass --reset-values and your
file, not --reuse-values: reused values carry every key the previous release accepted,
and v0.8.0 refuses the retired core.enabled and global.amfAddr, and
racora-core.mcc/mnc that disagree with global.network.plmn. my-values.yaml is the
file that carries your placement and identity; the registry stays a --set
(Install onto an Existing Kubernetes Cluster).
For an external core put global.core.provider and global.core.amfAddr in that file: the
overlay's amfAddr is a placeholder the chart refuses.
With the installer: INSTALL_RACORA_PLATFORM=kubernetes INSTALL_RACORA_VERSION=vX.Y.Z RACORA_HELM_EXTRA_ARGS='-f my-values.yaml', with the same INSTALL_RACORA_CORE,
INSTALL_RACORA_AMF_ADDR and RACORA_REGISTRY as at install. This platform keeps no
record, so the installer reads the deployed release's provider and refuses a run without
INSTALL_RACORA_CORE that would switch it: for an external core pass
INSTALL_RACORA_CORE=external INSTALL_RACORA_AMF_ADDR=<amf> on every installer run. The
installer always resets values, so the file is passed on every run.
Offline (k3s Platform)
Fetch the new release's bundle on a connected machine (`scripts/fetch-bundle.sh vX.Y.Z
Upgrading a v0.7.x Install to v0.8.0
v0.8.0 makes the core a provider, declares the network identity once, and adds the
Subscriber API with its own controller. Re-running the installer on the control node does
everything for an Open5GS install; what changes:
- The PLMN moves. The chart no longer has
racora-core.mccandmnc; the identity isglobal.network(plmn,tacs,slices), default PLMN90170, TAC 7, SST 1. AHelmChartConfigthat still keys the PLMN underracora-corewithoutglobal.networkis refused by the installer's pre-flight before the payload is staged or k3s restarted (an offline bundle's k3s binary is already copied into place by then), with the migration in the message: addglobal.networknext to the old keys first (v0.7.x ignores it; the new chart accepts the pair while both agree), upgrade, then drop theracora-corekey, or the whole object if the default is your identity. A pair that disagrees fails the render. - An external core declared the v0.7.x way (
core.enabled: falseandglobal.amfAddrin theHelmChartConfig) has noRACORA_COREin its install record, so a plain re-run would select Open5GS; the pre-flight refuses that. Re-run withINSTALL_RACORA_CORE=external INSTALL_RACORA_AMF_ADDR=<amf>(recorded from then on), then dropcoreandglobal.amfAddrfrom theHelmChartConfig: the v0.8.0 render is refused while they are there and v0.7.x keeps serving, and the upgrade completes once they are gone. Never drop them before the re-run, or v0.7.x deploys Open5GS. - The core rolls once, because its chart is renamed (
racora-coretoracora-core-open5gs). Its Deployment, Service and subscriber claim are adopted in place, same objects, same data, and the CU-CP re-associates NGAP by itself within seconds. Later chart bumps no longer roll it. - Values keys.
core.enabledandglobal.amfAddrbecomeglobal.core.provider(open5gsorexternal) andglobal.core.amfAddr;racora-core.*becomesracora-core-open5gs.*; the Open5GS chart'ssubscriber.sqn, which fed the image's QoS class field, issubscriber.qci.core.enabledandglobal.amfAddrfail the render with their replacement,racora-core.mcc/mncare accepted while they agree withglobal.network.plmn, and the other retired keys are ignored, so move them. - On a cluster you run the same notes apply through
helm upgrade: drop the retired keys from your values file first, then upgrade with--reset-valuesand the file. - New in
racora-system: theracora-core-controllerDeployment and thesubscribers.racora.ioCRD; the CRD asset is oneracora-crds.yaml. The old stagednrcell-crd.yamlis removed from the manifests directory; itsAddonobject is left in place on purpose. - Host units move from
core-support/to the provider module (cores/open5gs/host/), are re-rendered from there and recorded in/etc/racora/install.env, so the uninstall removes exactly those. - Subscribers already in the core (seeded, or added through the WebUI) are untouched.
Declaring one as a
Subscriberadopts it (Add Subscribers).
What Rolls, and What It Keeps
| Component | State | On an upgrade |
|---|---|---|
| racora-controller | stateless; identity, PCI and neighbors are written to the NRCell objects | loses nothing |
| racora-core-controller | stateless; every Subscriber is re-applied from the API on a cadence | loses nothing |
| NRCell and Subscriber objects | etcd (CRDs) | durable |
| CU-CP, CU-UP, CU-IP, DUs | stateless; configuration regenerated from the NRCells | roll |
| Device plugins | stateless DaemonSets | roll |
| Open5GS subscriber database | mongodb on the open5gs-mongodb PersistentVolumeClaim (persistence.enabled, default on, helm.sh/resource-policy: keep) | survives rolls and upgrades |
| ClickHouse, Tempo | emptyDir by default; a claim with racora-monitoring.<component>.storage.persistent=true | stored telemetry is lost on a roll |
| Grafana | emptyDir | dashboards you built by hand are lost; provisioned ones return |
The subscriber claim is ReadWriteOnce, which fits: the core is a Recreate singleton
pinned to the control node, and with k3s's local-path StorageClass the volume is
node-local. What an uninstall does with the claim is on Uninstall Racora; to wipe
the database alone, kubectl -n 5g-core delete pvc open5gs-mongodb. Declared Subscribers are re-applied by
the core controller, so a wiped database converges back to what is declared.
The CRDs have one version, v1alpha1, and no conversion. Re-applying an older
racora-crds.yaml over a schema the cluster already advanced can be rejected.
Two things to keep in mind:
- A release pins its controller and CU-IP together. If you stage components by hand, move the controller first: an older controller applies decisions from a newer CU-IP without the semantics for them.
- Old release images stay on the nodes until kubelet's image garbage collection runs,
which starts at its disk-usage threshold. Once the upgraded pods are verified,
sudo k3s ctr images rm <ref>prunes superseded pins. A hot-bumped image tag in theHelmChartConfigstill applies after the upgrade; remove that key from itsvaluesContentonce the release carries the pin you wanted, and keep the object, which also carries your identity and registry (Override Chart Values).
If an Upgrade Fails
There is no automatic rollback, and the two platforms fail differently.
- A cluster you run (
helm upgrade): the release is leftfailedorpending-upgrade.helm -n default history racorashows it andhelm -n default rollback racora <last-good-revision>restores the previous revision. - k3s (helm-controller): the
HelmChartsets nofailurePolicy, so helm-controller's default,reinstall, applies. When its job finds the releasefailedor stuckpending-*, it uninstalls and reinstalls it rather than leaving it wedged: every workload is recreated and the helm history is lost. The NRCells survive (the CRDs are applied out-of-band, and the namespaces and the subscriber claim carryhelm.sh/resource-policy: keep); telemetry on emptyDir does not. A render refusal (a retired key, an identity the core cannot serve) fails the job before anything is changed and leaves the previous release deployed.
To recover:
- Inspect. k3s:
kubectl -n kube-system get helmchart racoraandkubectl -n kube-system logs -l helmcharts.helm.cattle.io/chart=racora. Any platform:helm -n default history racoraand the crash-looping workload's logs (Troubleshoot). - Roll the images back by re-installing the previous release
(
INSTALL_RACORA_VERSION=<older>), or hot-pin the one bad image with aHelmChartConfigor your values file, and remove that key once the next release ships. Rolling images back is deterministic: every tag is immutable.