Autonomy has stopped being a slideware ambition and started producing audited numbers. In February 2026, an operator received the world’s first Level 4 autonomy validation from TM Forum for RAN energy optimization running on a live commercial network — cutting radio access energy consumption by roughly 20% with zero impact on customer experience KPIs (Rakuten Mobile, 2026). Months earlier, a European operator had achieved the first Level 4 certification under TM Forum’s formal ANLAV scheme for predictive cell energy management, in production since May 2024 (Converge Digest, 2025). A major Asian operator reports that its network operations center reached Level 4 with an 80% reduction in major network faults (TM Forum, 2025).
The financial stake is quantified, not rhetorical. STL Partners estimates that driving higher levels of network autonomy is worth approximately US$794 million per year for an average communications service provider (CSP) — around US$300 million in capex savings, US$350 million in opex savings and US$144 million in incremental revenue, equivalent to roughly 5% of revenue for a reference operator with US$15.6 billion in turnover, 31 million mobile and 12 million fixed subscribers (STL Partners). That is not the value of a tool purchase; it is the value of changing how the network is operated.
Yet the industry-wide picture remains sober. TM Forum’s own survey work finds only 21% of respondents operating at Level 3 or above overall — up from 19% a year earlier — across 125 respondents in 80 companies (TM Forum Research, 2025). Benchmark data reported in 2026 places about 31% of operators at Level 2, 17% at Level 3 and just 4% at Level 4 (RCR Wireless, 2026). The gap between the certified frontier and the average operating reality is the space this analysis maps: what the framework actually says, what the leaders have measured, and — most importantly — the concrete steps and AI decisions that take an operator from basic automation to genuine autonomy.
The thesis of this analysis: autonomous networks are not an AI procurement exercise. They are an operating-model transformation measured scenario by scenario, in which AI is introduced in disciplined waves — analytics first, copilots second, agents last — on top of a data foundation and inside a governance perimeter. Operators that follow that sequence are already banking energy, fault-reduction and workforce results. Operators that buy the top of the stack without the bottom will join the long history of stalled telco automation programs.
1. Why the autonomy journey has become unavoidable
The cost of standing still
Network complexity is compounding faster than operations headcount can scale. A modern operator runs 4G and 5G radio layers in parallel with legacy fixed access, an IP and optical transport fabric, cloud-native core functions distributed across national and edge data centers, and an OSS estate accumulated over two decades of vendor generations. Every added layer multiplies the alarm volume, the configuration surface and the cross-domain failure modes that a human-centric network operations center (NOC) must absorb. The traditional response — more tiers, more runbooks, more people — stopped scaling years ago.
Three pressures converge on the same conclusion:
- Cost structure: network operations and energy dominate the opex line. Energy alone is commonly estimated at 20–40% of network operating cost (industry estimates), and the radio access network is its largest consumer — which is precisely why the first certified Level 4 scenarios are energy scenarios.
- Service expectations: enterprise slices, private networks and latency-sensitive services carry contractual SLAs that manual, ticket-driven assurance cannot honor. Detecting, diagnosing and repairing within minutes requires machine-speed closed loops.
- Workforce reality: the engineers who know the legacy estate are retiring, and the profiles who could replace them do not want careers built on alarm triage. Autonomy is as much a talent strategy as a cost strategy.
The value on the table
The economics of autonomy have been modeled with unusual rigor for this industry. The STL Partners assessment referenced above decomposes the US$794 million annual figure into three mechanisms, each tied to specific operational capabilities rather than generic “AI benefits” (STL Partners):
- Capex efficiency (~US$300M/year): resource management and network planning driven by predictive models — investing where saturation will actually occur instead of where static utilization reports suggest, and sweating existing assets through optimization rather than overprovisioning.
- Opex reduction (~US$350M/year): automated fault management, energy optimization, field-force reduction through remote and predictive resolution, and the compression of manual change and provisioning work.
- Revenue uplift (~US$144M/year): faster time-to-market for services, SLA-backed premium connectivity that only autonomous assurance can guarantee, and improved retention from fewer customer-impacting incidents.
The distribution of that value matters for prioritization. The largest pools sit in resource operations — RAN energy, fault management, capacity planning — not in exotic new services. This is consistent with what the certified Level 4 cases actually did first: they went after energy and assurance, the domains where value is large, measurable and achievable with today’s AI. Operators designing their journey around the money, rather than around technology novelty, converge on the same short list of starting scenarios.
Automation projects versus autonomy as an operating model
Most operators already “do automation” — and that is precisely the trap. Scripted provisioning, rule-based ticket routing and scheduled health checks are Level 1–2 capabilities: they execute predefined actions faster, but a human still performs the awareness, analysis and decision work. Autonomy is a different construct: the system itself perceives state, diagnoses causes, decides within policy boundaries and executes — in a closed loop, continuously, with humans setting intent and auditing outcomes rather than approving each action.
The practical difference shows up in how work is funded and measured. Automation initiatives are typically project-funded: build a script, close the ticket backlog, declare success. Autonomy requires two parallel investment tracks — long-lead enablers (data platforms, governance frameworks, architectural standards, new teams) measured by adoption and stability, and short-lead use cases measured by capex/opex/revenue impact. Treating structural enablers like quick-win use cases leads to underfunding; treating quick wins like heavy initiatives leads to slow delivery. Every stalled transformation program exhibits one of those two confusions.
2. The TM Forum framework: a common language for autonomy
The industry converged on TM Forum’s Autonomous Networks framework because it solved a vocabulary problem. Before it, every operator and supplier graded their own homework — “self-optimizing”, “zero-touch” and “intelligent” meant whatever the slide author needed. The framework, developed through documents such as IG1218 (business requirements), IG1230/IG1251 (technical architecture) and IG1252 (level evaluation methodology), established a shared six-level scale and, critically, a standardized way to measure position on it (TM Forum IG1252).

The six levels, from L0 to L5
The scale describes who — human or system — performs each operational task:
- Level 0 — Manual operations: humans execute the full lifecycle; systems provide monitoring and alarms only.
- Level 1 — Assisted operations: the system executes pre-configured repetitive tasks (scripts, macros); humans initiate and control every cycle.
- Level 2 — Partial autonomy: closed-loop automation under static rules for specific units; the system acts, but humans still perform awareness, analysis and decision.
- Level 3 — Conditional autonomy: the system senses environment changes in real time and self-optimizes against dynamic policies; humans validate or approve the significant decisions.
- Level 4 — High autonomy: the structural jump. The system performs analysis and makes decisions itself, using predictive intelligence and continuous learning, across a defined multi-domain scope. Intent replaces instruction; humans supervise outcomes rather than approve actions.
- Level 5 — Full autonomy: the system generates its own operational intents and adapts across all domains and the entire lifecycle. No operator claims this today; it functions as the asymptote, not the plan.
The jump that matters commercially is Level 3 to Level 4. Up to Level 3, a human remains inside the decision loop, which caps both the speed and the scale of what automation can deliver — every action queue drains at human pace. At Level 4, the human moves from in the loop to on the loop: defining objectives and constraints, auditing behavior, and intervening by exception. That is what allows one engineer to supervise a portfolio of closed loops instead of a queue of tickets.
Five task dimensions, evaluated flow by flow
A level is not a network-wide badge; it is scored per operational flow across five task dimensions. The IG1252 methodology decomposes any operations process — RAN fault management, IP provisioning, core change management — into its sub-processes and asks, for each of five cognitive tasks, whether the human or the system performs it (TM Forum IG1252):
- Intent: who translates business objectives into operational goals and constraints.
- Awareness: who perceives network state — data collection, monitoring, prediction of emerging conditions.
- Analysis: who identifies the fault, performs root cause analysis (RCA) and generates candidate solutions.
- Decision: who evaluates the candidates and selects the action.
- Execution: who implements the action on the network.
This granularity is what makes the framework operational rather than decorative. An operator can honestly hold Level 3.4 in RAN outage handling, Level 2 in transport change management and Level 1 in enterprise service fulfillment simultaneously. The score profile tells the transformation team exactly which cognitive task in which flow is the bottleneck — typically analysis and decision, since awareness (telemetry) and execution (orchestrators) are the easier engineering problems. The AI investment plan falls directly out of that diagnosis.
How progress is measured and certified
The framework comes with an audit trail, which distinguishes it from marketing scales. Three instruments matter:
- ANLET (Autonomous Network Level Evaluation Tools): standardized, questionnaire-based tools (the GB1059 series) that decompose scenarios into tasks and produce a defensible level score, designed to expose weaknesses quickly (TM Forum IG1392).
- ANLAV (AN Level Assessment and Validation): TM Forum’s independent validation that an assessment was performed correctly — the scheme under which the first Level 4 certifications were issued in 2025–2026.
- The KBI/KEI/KCI hierarchy: Key Business Indicators (opex reduction, revenue), Key Effectiveness Indicators (MTTR, SLA compliance rate, energy-saving ratio) and Key Capability Indicators (automation rates per task). The hierarchy forces every technical capability to trace upward to a business number.
Two effectiveness metrics deserve a permanent place on the transformation dashboard. The operations automation rate — tasks completed with zero human involvement divided by total tasks — is the cleanest single proxy for level progression, and it approaches 100% for a scenario only at Level 4. Mean time to repair (MTTR), measured end-to-end from detection to verified restoration, is the metric where autonomy’s compounding effect is most visible: AI-driven event handling has publicly demonstrated compressions from hours to minutes, as the cases below show.
Where the industry actually stands
The honest reading of the adoption data is “moving, but from a low base”. TM Forum’s survey across 80 companies puts 21% of respondents at Level 3 or above network-wide, versus 19% a year earlier (TM Forum Research). Meanwhile the certified frontier is advancing much faster than the average: TM Forum reports describe the industry reaching a point of “significant change” toward Level 4, with multiple operators validating Level 4 in specific scenarios during 2025 and 2026 (Fierce Network, 2025).
The distinction between a validated scenario and network-wide maturity is the most abused nuance in this market. A Level 4 certificate covers one high-value scenario — RAN energy optimization, IP fault management — not the operator’s entire estate. Reading a certification press release as “operator X now runs a Level 4 network” is exactly the misreading the evaluation methodology was designed to prevent. The leaders themselves are explicit that they operate a portfolio of levels, raising the floor domain by domain.
3. Level 4 in production: what the leaders have measured
Four public cases define the current frontier, and their numbers are worth reading precisely. Each followed the same pattern: one high-value scenario, a closed loop with predictive AI in the analysis and decision tasks, and independent measurement of the result.

RAN energy autonomy at national scale
The first Level 4 validation on a live Open RAN network targeted energy, the largest controllable cost in the radio domain. Rakuten Mobile’s autonomous energy efficiency solution was validated by TM Forum at Level 4 for the RAN Energy Efficiency Optimization scenario (GB1059H) in February 2026 — a world first on a commercial network carrying production traffic. The system autonomously decides when and where to place cells into energy-saving states, delivering approximately 20% RAN energy conservation with zero measured impact on customer experience KPIs (Rakuten Mobile, 2026). Analyst commentary framed the validation as the moment Level 4 moved “from vision to early reality” (Dell’Oro Group, 2026).
Predictive energy management, formally certified
The first certification under the formal ANLAV scheme went to a Danish operator running predictive cell energy management in production since May 2024. TDC NET’s certified solution predicts per-cell traffic from historical signatures and load profiles, activates deep-sleep states in sub-second windows and restores capacity instantly when demand returns. Measured outcome: roughly 5% less energy per gigabyte transmitted across the targeted segment, an estimated 800 MWh saved and about 135 metric tons of CO2e avoided in 2024 (Converge Digest, 2025). The percentages look modest until multiplied by a national grid’s electricity prices — and until noted that they were achieved with no capacity sacrifice.
Agentic AI in the RAN: from hours to a minute
The most instructive agentic case in Europe pairs a large operator with a hyperscaler’s AI stack. Deutsche Telekom’s RAN Guardian agent — built with Google Cloud and launched into live operations in November 2025 — monitors radio performance continuously, classifies degradations, and triggers remediation autonomously. In its first month it executed over 100 autonomous remediation actions around high-traffic seasonal events; for 2026 the system has identified 237,000 network events, and it has reduced the time to manage major events from hours to around one minute — a more than 95% improvement (Deutsche Telekom, 2026). The program has since expanded into MINDR, a multi-agent framework for predictive diagnosis, signaling that the operator considers single-purpose agents merely the first step (Deutsche Telekom, 2025).
The Level 4 network operations center
The most complete NOC transformation on public record comes from China Mobile. Using a large telecom-tuned AI model exceeding 10 billion parameters, closed-loop tool chains and two agent classes — role-based copilots for staff and scenario-based agents for fault and complaint handling — the operator raised its NOC autonomy score from 3.2 to Level 4 under TM Forum assessment. Reported results: an 80% reduction in major network faults, more than 3,200 person-years of manual labor saved, over 30% reduction in backend O&M manpower and roughly 30% lower mean time to repair for faults and complaints (TM Forum case study, 2025). The operator’s stated direction is a “lights-out” operations model in which the default is machine handling and human touch is the exception.
What the frontier cases share
Strip the logos and the four cases are structurally identical:
- One scenario, not everything: each targeted a single high-value flow (energy, RAN assurance, NOC fault handling) rather than attempting network-wide autonomy.
- Predictive, not reactive AI: the differentiating capability in every case is prediction — of traffic, of degradation, of failure — placed in the analysis and decision tasks.
- Closed loop with guardrails: autonomous execution bounded by explicit policies, with customer-experience KPIs monitored as the veto condition.
- Independent measurement: results validated externally (TM Forum validation or certification) or published with auditable numbers — which is also what makes them internally defensible to CFOs.
4. The AI stack that gets you there
AI enters the autonomy journey in layers, and the layers are not interchangeable. The recurring failure mode of telco AI programs is buying capability at the top of the stack — copilots, agents — while the bottom — data — remains fragmented. The stack below is the architecture that the production cases have in common.

Layer 1: the unified data foundation
No model outperforms its telemetry. Autonomy requires real-time, normalized, cross-domain data: alarms, performance counters, configuration state, inventory, topology and service context accessible from one logical layer rather than extracted nightly from a dozen element managers. The architectural direction the industry has settled on is a unified data fabric — data products exposed via APIs over streaming pipelines, so that AI functions access live state where it resides instead of working on stale copies.
The component that separates leaders from laggards here is the knowledge layer. A topology-aware knowledge graph — encoding physical and logical dependencies between cells, links, network functions and services — is what allows an AI system to reason that five hundred alarms are one fiber cut, or that a proposed configuration change would isolate a critical site. In the agentic era this graph doubles as the grounding source that keeps language models anchored to network reality rather than plausible-sounding fiction. Operators consistently report that data unification and inventory accuracy, not model selection, consumed the majority of their transformation effort.
Layer 2: AIOps and machine learning
Classic machine learning does the industrial heavy lifting of autonomy, and it is unglamorous by design. This layer replaces static thresholds and rule cascades with learned behavior:
- Anomaly detection: unsupervised models learn each cell’s, link’s and function’s normal patterns and flag deviations — catching the “silent” quality degradations that never breach a static threshold.
- Event correlation and RCA: algorithmic compression of alarm storms into root-cause incidents using topology and causal inference, routinely cutting actionable event volume by an order of magnitude.
- Prediction: traffic forecasting for energy states, degradation prediction for proactive maintenance, capacity forecasting for planning — the capability that distinguishes every certified Level 4 case.
The maturity test for this layer is trust in production, not accuracy in the lab. A model that suppresses 90% of alarm noise is only valuable once the NOC actually stops looking at the raw feed. That trust is earned through months of shadow-mode operation, measured false-positive rates and explainable outputs — a sociotechnical process that no procurement cycle can shortcut.
Layer 3: generative AI copilots
Copilots change the interface to operations before they change operations itself. Role-based generative AI assistants let engineers query network state in natural language, summarize incident context across systems, translate vendor documentation into diagnostic steps and draft remediation playbooks with success probabilities based on historical outcomes. The measured value is cycle-time compression in the human-driven parts of the process — faster triage, faster diagnosis, faster handover between tiers.
Strategically, copilots are the trust bridge to agency. They expose AI reasoning to hundreds of engineers daily while keeping humans in full control of execution — building exactly the organizational confidence that Level 4 will later require. Operators that skipped this stage and went straight to autonomous execution have consistently faced internal resistance that stalled deployment longer than any technical gap.
Layer 4: agentic AI and intent-driven closed loops
Agents are goal-directed systems that plan, act through APIs and verify outcomes — the execution engine of Level 4. Where a copilot advises a human, a scenario-based agent owns a flow end to end: detect the degradation, diagnose it against the knowledge graph, select a remediation within policy, execute it through the orchestrator, confirm restoration, document the case. The operator-reported results above — 237,000 events identified, event management in about a minute, 80% fewer major faults — are what this architecture delivers when the layers beneath it are solid.
Two standards questions dominate agentic architecture, and both have industry answers forming:
- How agents receive goals: intent-based management, standardized in TM Forum’s TMF921 Intent Management API, which formalizes how business objectives are expressed to, negotiated with and reported on by autonomous domains (TMF921 v5.0).
- How agents talk to tools and each other: emerging open protocols — Model Context Protocol (MCP) for connecting models to network data and tools, and agent-to-agent protocols for delegation and conflict negotiation between agents from different suppliers — an area STL Partners identifies as decisive for avoiding a new generation of vendor lock-in (STL Partners, multi-agent systems).
Multi-agent conflict is the design problem nobody should discover in production. An energy agent that wants to sleep a cell and an experience agent that needs its capacity for a priority service will eventually disagree. Mature architectures resolve this with a supervisory layer — policy-as-code arbitration over all agents, with priorities set by business intent — rather than by hoping the conflict never occurs.
The trust layer: twins, grounding and human oversight
Probabilistic models acting on deterministic critical infrastructure require engineered safety, not optimism. The production deployments converge on three mechanisms:
- Digital twin validation: proposed changes are simulated against a live model of the network before touching production — the sandbox in which agent behavior is proven, misconfigurations are caught and blast radius is quantified.
- Knowledge grounding: agent plans are checked against the ontology and an allowed-procedure catalog before execution, blocking actions that violate engineering rules — the structural defense against model hallucination in network operations.
- Human-on-the-loop governance: autonomy budgets per scenario, mandatory human approval above defined risk thresholds, complete decision audit trails, and explainability requirements — no black-box actions on the live network.
5. The transformation playbook: five steps that generalize
The journeys that work follow a recognizable sequence, regardless of operator size or region. What follows is the generic playbook distilled from the certified cases, TM Forum’s assessment methodology and the published post-mortems of programs that stalled. It is deliberately vendor-neutral: every step names capabilities, not products.

Step 1: Assess the baseline — honestly and per flow
Every credible journey starts with a scored, per-domain, per-flow maturity baseline. Using the standardized evaluation methodology and tools (IG1252, the GB1059 ANLET questionnaires), the operator decomposes its operations — fault, change, provisioning, optimization, across RAN, transport, IP, core and cloud domains — and scores who performs intent, awareness, analysis, decision and execution in each flow. The output is a maturity heat map, not a single number.
Three elements make the assessment decision-grade rather than ceremonial:
- Baseline the KPIs simultaneously: capture current MTTR, automation rates, energy per gigabyte, truck rolls and SLA performance per flow — without the “before”, no “after” will ever be provable.
- Audit the data estate in the same pass: inventory accuracy, telemetry coverage and latency, alarm quality, data governance. AI readiness is set here, and the audit invariably finds the real blocker (fragmented inventory is the most common single finding).
- Map in-flight initiatives: most operators discover dozens of disconnected automation efforts; visibility of all of them is a prerequisite for a coherent roadmap rather than a portfolio of duplicates.
Step 2: Define the target state and choose high-value scenarios
The target is a portfolio of scenario-level ambitions with dates, not “Level 4 by 2028”. The selection discipline that the successful cases applied is consistent: rank candidate scenarios by value density (opex, energy or SLA exposure), technical feasibility with current data, and risk containment (can the blast radius be bounded and the action reversed). Energy optimization and fault management dominate the top of every honest ranking — high value, measurable, reversible — which is exactly where the certified Level 4 validations happened.
Anchor the target to business indicators from day one. Each selected scenario carries its KBI/KEI/KCI set: the business number it moves (energy cost, opex per subscriber), the effectiveness metrics that prove it (MTTR, automation rate, energy per GB) and the capability metrics that track buildout. This is also the moment to structure the two funding tracks — long-lead enablers versus short-lead use cases — so that the data platform is not asked to justify itself with a three-month payback.
Step 3: Build the data and platform foundation before buying AI
This is the step operators are most tempted to skip, and the step that decides the program. The foundation work is unglamorous: unify telemetry into a streaming data layer, fix inventory until the network model matches the network, build the topology knowledge graph, expose data as governed products through APIs, and standardize integration through open interfaces so that closed loops can span domains. Platform-centric beats project-centric: isolated AI experiments each rebuilding their own data pipeline is the documented anti-pattern that prevents scale.
Architectural choices made here determine the ceiling later:
- Open, standardized APIs between layers (TM Forum Open APIs, intent interfaces) so that domain autonomy can be composed into end-to-end autonomy instead of welded shut per vendor.
- A hierarchical loop architecture: fast local loops inside each domain (RAN, transport, core) with cross-domain orchestration above them — pushing intelligence down for speed while keeping end-to-end coherence above.
- The digital twin as first-class infrastructure, not a demo: every future autonomous action will need a place to be rehearsed.
Step 4: Introduce AI in three waves, each with guardrails
The AI sequencing that works is analytics, then copilots, then agents — value at every wave, trust compounding across them.
- Wave 1 — AIOps (typically months 1–9): anomaly detection, alarm correlation and prediction on the unified data layer. Target outcomes: order-of-magnitude alarm compression, first predictive maintenance saves, NOC attention shifted from noise to incidents. Runs in shadow mode until false-positive rates earn production trust.
- Wave 2 — Copilots and conditional autonomy (months 6–18): generative assistants for triage, diagnostics and change preparation; first closed loops move to Level 3 operation where the system proposes and executes after human approval. Intent interfaces are introduced so objectives, not instructions, start flowing downward.
- Wave 3 — Agents and Level 4 scenarios (months 12–36): scenario agents take end-to-end ownership of the selected high-value flows — detect, diagnose, decide, execute, verify — inside autonomy budgets, with twin-validated actions and human oversight by exception. Independent validation (ANLAV) closes the loop by making the achievement auditable.
The guardrails are not a later phase; they ship with wave 1. Explainability requirements, decision logging, rollback paths, KPI veto conditions (customer experience metrics that automatically halt an automation) and risk-tiered approval thresholds are defined before the first model touches production. The operators that publicized Level 4 results all emphasize the same point: autonomy scaled because trust was engineered, not assumed.
Step 5: Transform the operating model and govern by metrics
Autonomy fails organizationally before it fails technically, so the operating model is a step, not an afterthought. The workforce shifts from executing operations to engineering and supervising the systems that execute them; the sections below detail the roles. Governance becomes metric-driven: quarterly re-scoring of scenario levels, automation-rate and MTTR trajectories reviewed at executive level, and value attribution per scenario feeding the reinvestment decision. The transformation is complete for a scenario when the human role in it is intent-setting and audit — and the journey continues scenario by scenario until that is the norm rather than the exception.
6. The organizational dimension: roles, skills, governance
From ticket handlers to automation engineers
The workforce math is the most sensitive and least discussed part of the journey. The public numbers are blunt: thousands of person-years of manual work removed, 30% backend O&M manpower reductions, one technician maintaining a substantially larger site footprint. The operators that navigated this without organizational rupture did so by converting roles rather than only removing them — retraining operations engineers into automation developers who build, tune and supervise the closed loops, often through low-code platforms that let domain experts encode their knowledge without becoming software engineers.
The role set of an autonomous operation
Four role families recur across the mature deployments:
- Network strategist / intent owner: translates business objectives into intents, policies and constraint sets that bound what the autonomous systems may do; owns the trade-off decisions (energy versus experience, cost versus resilience).
- Data and AI analyst: owns telemetry quality, feature pipelines and model performance; converts operational questions into models and monitors drift, bias and false-positive economics in production.
- Automation / orchestration engineer: designs and maintains the closed loops and their integrations across domains; owns the digital-twin test harness and the rollback machinery.
- AI interaction and governance engineer: the emerging profile — curates prompts, playbooks and grounding knowledge for copilots and agents, evaluates agent decisions, and operates the guardrail stack (policy as code, audit trails, autonomy budgets).
Governance: the AI center of excellence and its limits
A central AI function accelerates the journey only if it enables rather than owns. The pattern that works: a center of excellence that sets standards (model lifecycle, data contracts, guardrail requirements, agent registration), provides shared platform services, and certifies scenario teams — while the scenario teams themselves live in the operational domains, close to the flows they automate. Fully centralized AI teams become bottlenecks; fully federated ones rebuild the same pipeline five times. The governance forum above both — reviewing autonomy budgets, incident post-mortems involving AI decisions, and the scenario-level scorecard — is where accountability for machine decisions formally lands.
7. Risks, limits and honest caveats
The scenario-versus-network gap
The headline risk is believing the frontier is the average. Level 4 exists in production — for specific scenarios, at operators that spent years on foundations. Network-wide, the industry median remains between Levels 2 and 3, and TM Forum’s own survey shows the overall needle moving two percentage points a year (TM Forum Research). Boards should expect a portfolio of levels for years, plan multi-year enabler funding accordingly, and treat any proposal promising network-wide Level 4 on a two-year horizon as unserious.
Probabilistic AI on deterministic infrastructure
Language-model errors that are amusing in a chatbot are outages in a network. The mitigation stack — knowledge grounding, twin validation, bounded action catalogs, KPI veto conditions — exists precisely because raw generative models cannot be trusted with change execution. The discipline to prohibit ungrounded, unauditable AI actions on the live network is a governance decision, and it must survive commercial pressure to ship faster. Configuration error remains a leading cause of major outages industry-wide; autonomy done carelessly automates the error at machine speed, while autonomy done properly is the strongest defense against it.
Legacy estate and data debt
Brownfield reality is the tax every incumbent pays on this journey. Decades of siloed OSS, inconsistent inventories, undocumented integrations and domain-specific tooling mean the data foundation step routinely takes longer than every AI step combined. The pragmatic pattern is progressive encapsulation: wrap legacy systems behind APIs and data products for the priority scenarios first, rather than gating the program on a total OSS replacement that history suggests will run late. Operators should also expect the assessment to surface organizational data debt — owners, definitions, quality accountability — that no platform purchase resolves.
The change-management failure mode
The technology now outruns the organization in most stalled programs. Engineers who distrust model outputs re-verify everything manually, erasing the cycle-time gains; middle management measured on headcount resists conversion plans; and finance treats enabler investment as discretionary the first difficult quarter. The countermeasures are known — shadow-mode trust building, explainability as a requirement, retraining commitments made early and publicly, and executive sponsorship that survives budget cycles — but they must be planned with the same rigor as the architecture. Culture is a dimension of the maturity model for a reason.
8. Conclusion: positions worth taking
For large operators, the frontier cases remove the excuse of impossibility. Level 4 is validated, in production, with audited energy, fault and workforce numbers attached. The strategic question is no longer whether but sequencing: baseline now, pick two or three high-value scenarios (energy and assurance lead every honest ranking), fund the data foundation as a multi-year enabler, and introduce AI in the analytics-copilots-agents order with guardrails from day one.
For mid-sized operators, the framework is the equalizer. The standardized methodology, evaluation tools and open APIs mean a focused operator can reach Level 4 in one scenario without hyperscaler-class budgets — the certified cases include national operators of modest size. The discipline that matters is scope: one scenario proven and banked beats five scenarios perpetually in pilot.
For suppliers and integrators, openness is becoming the qualification criterion. Operators burned by closed automation stacks are writing intent interfaces, open APIs and agent interoperability into procurement. Products that participate in multi-vendor closed loops will be composed into autonomy architectures; products that only automate their own silo will be confined to it.
The transferable principles compress into five lines:
- Measure before you automate: a scored baseline and KPI capture precede every credible journey.
- Follow the value density: energy and assurance first; exotic scenarios after the foundation pays for itself.
- Data before models, models before agents: the stack is a dependency chain, not a menu.
- Engineer trust explicitly: twins, grounding, audit trails and human-on-the-loop governance are the price of autonomous execution.
- Transform the organization in parallel: roles, skills and metric-driven governance decide whether the technology’s results survive contact with the operating model.
The autonomous networks journey will define the cost structure and service capability of the next telco decade. The framework is standardized, the measurement is auditable, the first US$794 million-scale prize has been sized, and the leaders have published their receipts. What separates the operators that capture this from those that watch it is no longer information — it is the discipline to execute the sequence.
At TelcoCrux, we separate the AI use cases that move telco P&Ls from the ones that only move slide decks. If you want a pragmatic, vendor-neutral view of where your autonomy journey actually stands — and which scenarios would pay first, let’s talk.



