Skip to main content

Upgrade Racora

A Racora release is one chart with one pinned image set. Upgrading means applying the next release to the cluster you have: on the k3s platform by re-running the installer on the control node, on a cluster you run with helm upgrade. Kubernetes then rolls exactly the pods whose spec changed, on the nodes they run on. There is no RAN version per node: the RAN is the control node's release, and a worker only ever runs the pods scheduled onto it.

k3s Platform​

On the server only:

curl -sfL https://get.racora.io | INSTALL_RACORA_VERSION=vX.Y.Z sh -

The re-run keeps the core provider and AMF address recorded at install time (/etc/racora/install.env) unless you give others; any INSTALL_K3S_* or K3S_* flag you set at install must be given again, because k3s rewrites its service file from them. The installer downloads the release's k3s binary only if the pinned upstream version changed, re-renders the core provider's host units, restarts k3s when its binary or service file changed, and only then stages the payload: the CRDs are re-applied out-of-band (racora-crds.yaml) and the HelmChart is rewritten with the new chart package. Same name, so k3s's helm-controller runs helm upgrade, and the Deployments roll with the new images and configuration; a DU on a worker pulls its new image from the registry.

Workers usually need nothing. Re-run node mode on a worker (with its K3S_URL, K3S_TOKEN and the same INSTALL_RACORA_VERSION as the control node) only when the k3s pin moved, which the k3s column of Release History shows: on an already provisioned host it validates and joins straight away, swapping the agent binary. After that restart, the USRP device plugin on the node must re-register (Add a Radio Node).

A Cluster You Run​

A plain helm upgrade with the two overlays from the release's installer package (Install onto an Existing Kubernetes Cluster); no host is involved:

helm upgrade --install racora oci://registry.gitlab.com/cognitive-network-solutions/racora/charts/racora \
--namespace default --version X.Y.Z --reset-values \
-f installer/platforms/kubernetes/values.yaml \
-f installer/cores/<provider>/values.yaml \
-f my-values.yaml \
--set global.systemDefaultRegistry=<registry>

This is the platform's command with --reset-values and the new release's installer/ overlays. Pass --reset-values and your file, not --reuse-values: reused values carry every key the previous release accepted, and v0.8.0 refuses the retired core.enabled and global.amfAddr, and racora-core.mcc/mnc that disagree with global.network.plmn. my-values.yaml is the file that carries your placement and identity; the registry stays a --set (Install onto an Existing Kubernetes Cluster). For an external core put global.core.provider and global.core.amfAddr in that file: the overlay's amfAddr is a placeholder the chart refuses.

With the installer: INSTALL_RACORA_PLATFORM=kubernetes INSTALL_RACORA_VERSION=vX.Y.Z RACORA_HELM_EXTRA_ARGS='-f my-values.yaml', with the same INSTALL_RACORA_CORE, INSTALL_RACORA_AMF_ADDR and RACORA_REGISTRY as at install. This platform keeps no record, so the installer reads the deployed release's provider and refuses a run without INSTALL_RACORA_CORE that would switch it: for an external core pass INSTALL_RACORA_CORE=external INSTALL_RACORA_AMF_ADDR=<amf> on every installer run. The installer always resets values, so the file is passed on every run.

Offline (k3s Platform)​

Fetch the new release's bundle on a connected machine (`scripts/fetch-bundle.sh vX.Y.Z

`, from the repository at the tag; a release whose page lists no image archive cannot be upgraded offline), carry it over, and re-run the installer on the control node with `INSTALL_RACORA_ARTIFACT_PATH=`. The image archive is imported into the control node's containerd only; nothing propagates images to workers. Re-run `node` mode on each worker with the same bundle path, or import the archive there yourself ([Use Your Own Registry or Install Offline](registries-and-offline.md)).

Upgrading a v0.7.x Install to v0.8.0​

v0.8.0 makes the core a provider, declares the network identity once, and adds the Subscriber API with its own controller. Re-running the installer on the control node does everything for an Open5GS install; what changes:

  • The PLMN moves. The chart no longer has racora-core.mcc and mnc; the identity is global.network (plmn, tacs, slices), default PLMN 90170, TAC 7, SST 1. A HelmChartConfig that still keys the PLMN under racora-core without global.network is refused by the installer's pre-flight before the payload is staged or k3s restarted (an offline bundle's k3s binary is already copied into place by then), with the migration in the message: add global.network next to the old keys first (v0.7.x ignores it; the new chart accepts the pair while both agree), upgrade, then drop the racora-core key, or the whole object if the default is your identity. A pair that disagrees fails the render.
  • An external core declared the v0.7.x way (core.enabled: false and global.amfAddr in the HelmChartConfig) has no RACORA_CORE in its install record, so a plain re-run would select Open5GS; the pre-flight refuses that. Re-run with INSTALL_RACORA_CORE=external INSTALL_RACORA_AMF_ADDR=<amf> (recorded from then on), then drop core and global.amfAddr from the HelmChartConfig: the v0.8.0 render is refused while they are there and v0.7.x keeps serving, and the upgrade completes once they are gone. Never drop them before the re-run, or v0.7.x deploys Open5GS.
  • The core rolls once, because its chart is renamed (racora-core to racora-core-open5gs). Its Deployment, Service and subscriber claim are adopted in place, same objects, same data, and the CU-CP re-associates NGAP by itself within seconds. Later chart bumps no longer roll it.
  • Values keys. core.enabled and global.amfAddr become global.core.provider (open5gs or external) and global.core.amfAddr; racora-core.* becomes racora-core-open5gs.*; the Open5GS chart's subscriber.sqn, which fed the image's QoS class field, is subscriber.qci. core.enabled and global.amfAddr fail the render with their replacement, racora-core.mcc/mnc are accepted while they agree with global.network.plmn, and the other retired keys are ignored, so move them.
  • On a cluster you run the same notes apply through helm upgrade: drop the retired keys from your values file first, then upgrade with --reset-values and the file.
  • New in racora-system: the racora-core-controller Deployment and the subscribers.racora.io CRD; the CRD asset is one racora-crds.yaml. The old staged nrcell-crd.yaml is removed from the manifests directory; its Addon object is left in place on purpose.
  • Host units move from core-support/ to the provider module (cores/open5gs/host/), are re-rendered from there and recorded in /etc/racora/install.env, so the uninstall removes exactly those.
  • Subscribers already in the core (seeded, or added through the WebUI) are untouched. Declaring one as a Subscriber adopts it (Add Subscribers).

What Rolls, and What It Keeps​

ComponentStateOn an upgrade
racora-controllerstateless; identity, PCI and neighbors are written to the NRCell objectsloses nothing
racora-core-controllerstateless; every Subscriber is re-applied from the API on a cadenceloses nothing
NRCell and Subscriber objectsetcd (CRDs)durable
CU-CP, CU-UP, CU-IP, DUsstateless; configuration regenerated from the NRCellsroll
Device pluginsstateless DaemonSetsroll
Open5GS subscriber databasemongodb on the open5gs-mongodb PersistentVolumeClaim (persistence.enabled, default on, helm.sh/resource-policy: keep)survives rolls and upgrades
ClickHouse, TempoemptyDir by default; a claim with racora-monitoring.<component>.storage.persistent=truestored telemetry is lost on a roll
GrafanaemptyDirdashboards you built by hand are lost; provisioned ones return

The subscriber claim is ReadWriteOnce, which fits: the core is a Recreate singleton pinned to the control node, and with k3s's local-path StorageClass the volume is node-local. What an uninstall does with the claim is on Uninstall Racora; to wipe the database alone, kubectl -n 5g-core delete pvc open5gs-mongodb. Declared Subscribers are re-applied by the core controller, so a wiped database converges back to what is declared.

The CRDs have one version, v1alpha1, and no conversion. Re-applying an older racora-crds.yaml over a schema the cluster already advanced can be rejected.

Two things to keep in mind:

  • A release pins its controller and CU-IP together. If you stage components by hand, move the controller first: an older controller applies decisions from a newer CU-IP without the semantics for them.
  • Old release images stay on the nodes until kubelet's image garbage collection runs, which starts at its disk-usage threshold. Once the upgraded pods are verified, sudo k3s ctr images rm <ref> prunes superseded pins. A hot-bumped image tag in the HelmChartConfig still applies after the upgrade; remove that key from its valuesContent once the release carries the pin you wanted, and keep the object, which also carries your identity and registry (Override Chart Values).

If an Upgrade Fails​

There is no automatic rollback, and the two platforms fail differently.

  • A cluster you run (helm upgrade): the release is left failed or pending-upgrade. helm -n default history racora shows it and helm -n default rollback racora <last-good-revision> restores the previous revision.
  • k3s (helm-controller): the HelmChart sets no failurePolicy, so helm-controller's default, reinstall, applies. When its job finds the release failed or stuck pending-*, it uninstalls and reinstalls it rather than leaving it wedged: every workload is recreated and the helm history is lost. The NRCells survive (the CRDs are applied out-of-band, and the namespaces and the subscriber claim carry helm.sh/resource-policy: keep); telemetry on emptyDir does not. A render refusal (a retired key, an identity the core cannot serve) fails the job before anything is changed and leaves the previous release deployed.

To recover:

  1. Inspect. k3s: kubectl -n kube-system get helmchart racora and kubectl -n kube-system logs -l helmcharts.helm.cattle.io/chart=racora. Any platform: helm -n default history racora and the crash-looping workload's logs (Troubleshoot).
  2. Roll the images back by re-installing the previous release (INSTALL_RACORA_VERSION=<older>), or hot-pin the one bad image with a HelmChartConfig or your values file, and remove that key once the next release ships. Rolling images back is deterministic: every tag is immutable.