Skip to main content

Write an Intelligence Function

An Intelligence Function reads the CU-IP substrate and emits decisions back toward the gNB. Two shipping functions are the worked references:

  • pci (ran_functions/pci/) — GPU-side, deterministic 2-hop message passing over the graph; the original pattern-setter.
  • anr (ran_functions/anr/) — CPU-side, stateful hysteresis over the graph's overlap edges; the reference for a factory-built stateful engine.

This guide is the general pattern: the contract, the framework pieces, and the scaffold checklist for function number three.

The Contract​

The data a function touches is defined in rann_core/rann_core/data/ — the wire interface between CU-IP and the rest of the gNB (the decision dataclass and the logging/tracing helpers live beside it in rann_core/, linked below):

  • Four substrate tables (output, substrate → your function): graph_nodes, graph_edges, reports, decisions.
  • Two ingest schemas (input, producers → substrate): ue_report_v2, cell_config_v2.

A build may extend any table with additional columns; the substrate passes unknown columns through unchanged to consumers, so a function can carry private features without changing the base contract.

A function's output is an Arrow table conforming to DECISION_SCHEMA, built from the decision contract in rann_core/rann_core/decisions.py: a Decision carries the target cell, the field to change (target_field), the recommended value, a reason, and a confidence. decision_id is deterministic ({function}:{target_name}:{target_field}) so consumers can dedup across polls. One rule of thumb from ANR: when target_field is a list (e.g. spec.neighbors), the recommended_value must be the complete desired list — JSON merge-patch replaces arrays wholesale, so a delta would erase the rest.

Served decisions are carried out of the substrate over RANN-D (/decisions/active on Arrow Flight) — your function never talks to CU-CP directly. The racora controller polls that endpoint and routes by function + target_field.

The Framework​

Registry and Env Gating​

In-process engines register in engine_registry.py: one AVAILABLE_ENGINES entry mapping the function name to a module:factory path.

  • The factory is called once at load with your function's slice of CUIP_ENGINE_CONFIG and returns the engine callable engine(ctx) -> pa.Table (DECISION_SCHEMA). Factories make stateful engines natural — ANR's hysteresis state, the PCI engine's warm solver — and keep config load-time, not batch-time.
  • CUIP_ENGINES (comma-separated names) selects which functions run; the default is pci,anr. Imports are lazy — a disabled function's dependencies are never imported — and unknown names fail the boot fast.
  • CUIP_ENGINE_CONFIG is one JSON object keyed by function name, e.g. '{"anr": {"theta_add": 0.4, "ownership": "report-only"}}'. Invalid JSON fails fast too: silently running on defaults when the operator thought they configured something is the quiet failure this framework refuses.

EngineContext​

Your engine receives one EngineContext per ingest batch, shared by all engines:

FieldWhat it is
nodes, edgesthe graph frames (cudf in production, pandas in tests — opaque to the framework)
configured_neighbors{cell: {'neighbors': [...], 'sources': {ref: 'manual'|'anr'}}} from the NRCell CRDs
nci_to_name / name_to_nciCRD identity maps
now_tsepoch seconds at context build — use this, never your own clock, so evaluation is reproducible in tests

Extension fields are added to the dataclass, never as new positional arguments — an engine written against an older context keeps working.

Compute Placement Is the Engine's Choice​

The framework hands over frames and doesn't care where you compute. PCI stays on GPU (torch tensor ops over the whole graph). ANR copies the tiny cells²-row tables to CPU at entry (graph_math.to_cpu_frame) — which makes the entire engine unit-testable in the GPU-free CI tier. Pick per function: bulk math on the graph wants the GPU; small branchy logic wants testability.

Evaluation, Isolation, Validation​

run_engines (decision_emit.py) runs each engine inside its own infer.eval span (attribute function), validates the output against DECISION_SCHEMA and rejects tables that fail (the warning names the offending function), and isolates failures — a crashing or wrong-schema engine costs only its own decisions, never the batch or the other engines.

Emit Hooks (Span Attributes)​

Every decision gets one rannd.emit span with generic attributes (decision.current, decision.recommended, identity, reason, confidence), and its trace_context is stamped with that span's W3C traceparent — consumers parent their dispatch spans to it, so decisions render as complete cross-repo traces in Tempo with zero tracing code in the engine. Add a per-function entry to EMIT_ATTR_HOOKS for function-shaped attributes: pci preserves the pci_* names existing dashboards filter on; anr adds add/remove deltas and the neighbor count. A hook failure never costs the emit. Span naming and attribute conventions are the Racora span schema: https://docs.racora.io/reference/observability/.

Logging​

Use rann_core.cns_logging.function_logger('<name>') — it prefixes every line with function=<name> on the INFER sublayer, so one grep isolates your function in the cuip logs.

Naming Convention​

  • Function name: short, lowercase, stable — it is the function value on every Decision, the CUIP_ENGINES token, the config key, and the span attribute.
  • Pure algorithm: rann_core/rann_core/<name>/ (stdlib/numpy only if you want GPU-free tests; this layer never imports dataframes).
  • Substrate adapter: substrate/rann_substrate_processing/<name>_engine/ exposing make_engine(config).
  • Docs: ran_functions/<name>/README.md (design, config reference, limitations).

What Never Changes​

Adding a function is function-local code plus declarative controller registrations, not a control-plane change. The wire contract (rann_core/decisions.py, DECISION_SCHEMA), the decision transport and the controller's dispatch routing stay as they are; what a new function may add controller-side is registration data — a post-apply policy, a status-patch builder, new precondition handlers. That is the point of separating what to change from when it is safe and from how to apply it.

Scaffold Checklist​

  1. Write the pure core in rann_core/rann_core/<name>/ with its unit tests in rann_core/tests/.
  2. Write the adapter package substrate/rann_substrate_processing/<name>_engine/ with make_engine(config); adapter tests in substrate/tests/.
  3. Register: add the AVAILABLE_ENGINES entry (append — order is evaluation order). Decide whether it joins _DEFAULT_ENGINES or stays opt-in via CUIP_ENGINES.
  4. Add an EMIT_ATTR_HOOKS entry if your function has attributes worth filtering on in Tempo.
  5. Add both test files to the pytest-gpu-free list in .gitlab-ci.yml (GPU-dependent parts go to the manual RAPIDS tier instead — see test_graph_builder_gpu.py for the shape).
  6. Write ran_functions/<name>/README.md.
  7. Controller side (racora repo): the dispatcher routes any new function + target_field generically, but check whether your function needs its own apply semantics or CRD status audit trail there before shipping the loop end to end.

Test Recipe​

  • Pure core: exhaustive, GPU-free, no dataframes (test_anr_hysteresis.py is the model — thresholds, sustain windows, state across calls).
  • Adapter: drive engine(ctx) with pandas frames and assert on the decoded Decisions (test_anr_engine.py); include one load_engines('<name>', {...}) test to prove the real registry path imports GPU-free (skip this if your adapter imports cudf/torch — the pci entry can only be exercised in the RAPIDS tier).
  • Determinism: same ctx in, same decisions out; sort everything you emit.
  • Use ctx.now_ts for time so tests control the clock.

Standalone Functions​

Out-of-process functions can read every substrate table over the Flight endpoints (FlightSubstrateClient in rann_core.data; wire contract in substrate/API.md), but decisions are written only by in-process functions: there is no decisions write path over Flight, so a function that actuates runs inside CU-IP.