Generated at build time from
charts/racora-monitoring/README.md(helm-docs:values.yaml+README.md.gotmpl); edit the source, not this page.
racora-monitoring
The observability stack for Racora. Four components, one wire contract — every producer sends OTLP to the collector, which fans out to ClickHouse (logs) and Tempo (traces); Grafana reads from both with pre-provisioned datasources.
What this installs
| Component | Image | Role |
|---|---|---|
| OTel Collector | mirrored-opentelemetry-collector-contrib:0.115.1 | Single OTLP ingress (gRPC :4317 + HTTP :4318), fans out logs→ClickHouse, traces→Tempo; its own metrics on :8888 (exposed, not scraped) |
| ClickHouse | mirrored-clickhouse-server:24.10 | Log SQL store. The otel_logs table is auto-created by the exporter on first connect (the exporter is on the logs pipeline only — no trace or metric tables) |
| Tempo | mirrored-tempo:2.6.1 | Distributed-trace store, single-binary monolithic mode; metrics-generator enabled with no remote-write target |
| Grafana | mirrored-grafana:11.4.0 | UI with Tempo + ClickHouse datasources and the dashboards under files/dashboards/ pre-provisioned (folder "Racora"), NodePort 30300 |
The images are the upstream ones, mirrored unmodified into the racora registry
under these mirrored-* names (versions.yaml → third_party); the distribution
re-roots them with global.systemDefaultRegistry.
All four are independently togglable (<component>.enabled defaults to true).
Requires
racora-base(monitoringnamespace)
No CRD dependency — this stack is RAN-agnostic and could be reused unchanged for any cluster.
Install
helm install racora-monitoring ./charts/racora-monitoring
Values
| Key | Type | Default | Description |
|---|---|---|---|
| namespace | string | "monitoring" | Namespace the monitoring stack deploys into. Must already exist (racora-base creates it). |
| clusterLabel | string | "racora" | Cluster identity attached to every record by the OTel collector. Useful when federating multiple gNB clusters into a single observability backend (each cluster's records are disambiguated by this label). Matches the OTel resource attribute set by the python services (cns_logging.configure_otel). |
| imageRegistry | string | "" | Image registry prefix. The distribution sets global.systemDefaultRegistry (the racora registry) and the third-party images below are mirrored into it under mirrored-* names — one re-rootable namespace, no Docker Hub pulls at runtime. Left empty here; the umbrella's global value drives it. |
| clickhouse.enabled | bool | true | Render ClickHouse. |
| clickhouse.image.repository | string | "mirrored-clickhouse-server" | ClickHouse image name. |
| clickhouse.image.tag | string | "24.10" | ClickHouse image tag; the upstream image, mirrored unmodified into the racora registry; pinned in versions.yaml (third_party). |
| clickhouse.image.pullPolicy | string | "IfNotPresent" | Image pull policy — IfNotPresent, because the images ship with the distribution. |
| clickhouse.db | string | "otel" | Database the collector's ClickHouse exporter writes otel_logs into. |
| clickhouse.user | string | "default" | ClickHouse user; no password — the store is in-cluster only. |
| clickhouse.resources | object | {"limits":{"cpu":2,"memory":"4Gi"},"requests":{"cpu":"200m","memory":"512Mi"}} | Pod resources for ClickHouse. |
| clickhouse.storage.persistent | bool | false | Use a PersistentVolumeClaim instead of emptyDir, so the data survives pod restarts. |
| clickhouse.storage.size | string | "50Gi" | Claim size when persistent. |
| clickhouse.storage.storageClassName | string | "" | StorageClass for the claim; empty = the cluster default. |
| tempo.enabled | bool | true | Render Tempo. |
| tempo.image.repository | string | "mirrored-tempo" | Tempo image name. |
| tempo.image.tag | string | "2.6.1" | Tempo image tag; the upstream image, mirrored unmodified into the racora registry; pinned in versions.yaml (third_party). |
| tempo.image.pullPolicy | string | "IfNotPresent" | Image pull policy — IfNotPresent, because the images ship with the distribution. |
| tempo.blockRetention | string | "24h" | Trace retention (Tempo compactor block_retention). |
| tempo.resources | object | {"limits":{"cpu":1,"memory":"2Gi"},"requests":{"cpu":"100m","memory":"256Mi"}} | Pod resources for Tempo. |
| tempo.storage.persistent | bool | false | Use a PersistentVolumeClaim instead of emptyDir, so the data survives pod restarts. |
| tempo.storage.size | string | "20Gi" | Claim size when persistent. |
| tempo.storage.storageClassName | string | "" | StorageClass for the claim; empty = the cluster default. |
| otelCollector.enabled | bool | true | Render OTel Collector. |
| otelCollector.image.repository | string | "mirrored-opentelemetry-collector-contrib" | OTel Collector image name. |
| otelCollector.image.tag | string | "0.115.1" | OTel Collector image tag; the upstream image, mirrored unmodified into the racora registry; pinned in versions.yaml (third_party). |
| otelCollector.image.pullPolicy | string | "IfNotPresent" | Image pull policy — IfNotPresent, because the images ship with the distribution. |
| otelCollector.resources | object | {"limits":{"cpu":1,"memory":"1Gi"},"requests":{"cpu":"100m","memory":"256Mi"}} | Pod resources for OTel Collector. |
| otelCollector.logsTtl | string | "24h" | Retention for logs in ClickHouse — independent of trace retention (which lives in tempo.blockRetention). |
| grafana.enabled | bool | true | Render Grafana. |
| grafana.image.repository | string | "mirrored-grafana" | Grafana image name. |
| grafana.image.tag | string | "11.4.0" | Grafana image tag; the upstream image, mirrored unmodified into the racora registry; pinned in versions.yaml (third_party). |
| grafana.image.pullPolicy | string | "IfNotPresent" | Image pull policy — IfNotPresent, because the images ship with the distribution. |
| grafana.service.type | string | "NodePort" | Service type: NodePort for browser access from outside the cluster, ClusterIP to port-forward instead. |
| grafana.service.nodePort | int | 30300 | NodePort used when service.type is NodePort. |
| grafana.anonymousAccess | bool | true | Anonymous access to the UI (no login) with the role below. Viewer can browse dashboards and query; Editor could also change them. Tighten (anonymousAccess: false) when exposing Grafana beyond a lab. |
| grafana.anonymousRole | string | "Viewer" | Role anonymous users get: Viewer (browse and query) or Editor (also change dashboards). |
| nodeSelector | object | {} | Node selector for the monitoring workloads (applied per workload unless overridden). |
| tolerations | list | [] | Tolerations for the monitoring workloads (applied per workload unless overridden). |
| affinity | object | {} | Affinity rules for the monitoring workloads (applied per workload unless overridden). |
Producer wiring
Every Racora/OCUDU service writes OTLP to the collector at
otel-collector.monitoring.svc.cluster.local. The collector accepts
both OTLP transports, and the producers split by SDK:
- gRPC
:4317— the Python components (controller, CU-IP) viacns_logging.configure_otel() - HTTP
:4318— the gNB/DU (its C++ OTel exporter posts to/v1/traces//v1/logs; the controller injects the endpoint into each DU Deployment)
Configured per-service via the OTEL_EXPORTER_OTLP_ENDPOINT env var.
The Python bootstrap soft-imports OTel packages so absent collector =
stdout logging, no spans, no failure.
Storage caveat
ClickHouse + Tempo both default to emptyDir. Pod restart wipes the
data. Fine for the iteration-mode reality where the system is reset
frequently. For real distribution use, flip <component>.storage.persistent: true
in values; the chart provisions a PVC automatically.
Verify
kubectl get pods -n monitoring
# Expected:
# clickhouse-xxx 1/1 Running
# tempo-xxx 1/1 Running
# otel-collector-xxx 1/1 Running
# grafana-xxx 1/1 Running
# Grafana (if NodePort):
curl -s http://localhost:30300/api/health
Once installed, the "Failed to export logs to otel-collector..." warnings in every other service's logs should immediately stop.