Four Layers, Four Phases, One Contract: The Development Model for an Agentic NOC/SOC

The technology decision is no longer the hard one. An operator that wants to run its network and security operations centre with AI agents can buy the models, the orchestration layer, the graph database and the observability stack from a mature supply chain, and the industry has agreed on a shared vocabulary for the destination through the TM Forum autonomous network levels and the ETSI zero-touch framework (TM Forum IG1252, ETSI ISG ZSM). What is not settled, in most operators, is far more basic: who decides what gets automated, who builds it, who is allowed to let a piece of software change a live network, and who answers for the euros when it does.

The evidence that the constraint is organizational rather than technical is now measurable. In NVIDIA’s fourth annual industry survey, published in February 2026 with roughly 1,000 telecom respondents, 90% reported that AI is both raising revenue and lowering cost, 89% planned to increase AI budgets in 2026 against 65% a year earlier, and autonomous networks ranked as the top return-on-investment use case at 50% — yet 88% of the same respondents still placed their own operation between Level 1 and Level 3 of the TM Forum autonomy scale (State of AI in Telecommunications 2026). Money, conviction and tooling are all present. Autonomy is not.

The gap between funded intent and operating reality has a specific shape, and it is worth naming before designing anything. Gartner’s June 2025 forecast that more than 40% of agentic AI projects will be cancelled by the end of 2027 attributes the failures to escalating cost, unclear business value and inadequate risk controls — three governance defects, not three engineering defects (Gartner). This article works from the global question down to the concrete one: how a telecom operator with a conventional legacy base should organize the working relationship between its network operations function, its systems integrator and its IT and systems area to build a NOC/SOC that runs on agents without losing control of the network.

The thesis: an AI-agent NOC/SOC is not delivered by a programme, it is delivered by an operating model. The operator that wins reassigns three things — the unit of work moves from the request to the operational scenario, the control point moves from a uniform pre-approval to a risk-tiered guardrail, and accountability splits cleanly into who owns the outcome, who owns the build and who owns the platform. Everything else, including model choice, is downstream of those three reassignments.

Almost every operator runs network automation through the same four-step chain, and it was a reasonable design. Operations identifies a need, raises a demand item, the systems or IT area reviews it for architecture, security and integration impact, a development team builds it, and the result is deployed to production. The chain assumes that the output is software, that the requester cannot build it, and that the risk of the change is roughly constant across requests. All three assumptions fail once the deliverable is an agent that reasons over operational context and can act on the network.

The first change is that the useful unit of work stops being a deliverable and becomes a process. A script that extracts a report is a deliverable: it has a specification, an acceptance test and an end date. An agent that triages access-network degradations is a process participant: its value depends on the quality of the surrounding workflow, its accuracy drifts as the network changes, and it needs an owner for as long as it runs. Managing the second thing with the governance built for the first produces automations that are accepted, deployed and then quietly abandoned.

The second change is that risk stops being uniform and becomes bimodal. In a portfolio of agentic use cases, a large share are read-only — summarising an incident, correlating alarms, drafting a root-cause narrative, answering a topology question in natural language — and carry roughly the risk profile of a dashboard. A much smaller share write: they open and close tickets, push configuration, change a radio parameter, isolate a port, block a prefix. Applying one review depth to both means the read-only majority waits behind the write minority, and the queue becomes the programme’s dominant cost.

The third change is that the requester can now build. Low-code agent builders, retrieval frameworks and connector catalogues put assembly within reach of a network engineer who understands the procedure better than any developer will. That is an opportunity and a governance problem at the same time: if the official route takes three months and the unofficial route takes an afternoon, the operation will take the afternoon, and the operator ends up with an inventory of undocumented automations touching production with a departed author and a shared service account.

Uniform gating does not fail loudly; it fails as four recognisable patterns that most transformation programmes exhibit simultaneously. Each of them is a predictable consequence of applying one control depth to a portfolio with a bimodal risk distribution:

  • Shadow automation. Scripts, notebooks and agent prototypes running outside the platform, without inventory, ownership, identity or logging. They work until they do not, and nobody can say what they touch.
  • The pilot graveyard. A portfolio of proofs of concept that each demonstrated feasibility, none of which was designed for reuse, and none of which cleared the production gate because the gate was designed for a different kind of change.
  • The integrator as ticket-taker. A partner with deep capability consumed as a body shop, delivering exactly what was specified in a demand item written by someone who could not see the whole process.
  • Systems as universal approver. An IT function that becomes the bottleneck for every initiative regardless of risk, is resented for it, and — because attention is finite — ends up reviewing the low-risk majority with the same superficiality as everything else.

The fourth pattern is the most expensive because it is self-defeating. A control function that is mandatory everywhere is thorough nowhere. The point of the operating model proposed here is not to reduce control; it is to concentrate control where the blast radius justifies it and to pre-approve the path everywhere else, which is the only way a scarce architecture and security capability can be spent on the changes that can actually take a network down.

Any honest operating model has to be designed for the estate the operator actually has, not the one in the reference architecture. A conventional European or Latin American operator arrives at this problem with a recognisable inheritance, and each element of it constrains what the first year of work can contain:

  • Multiple inventories that disagree. Physical, logical and service inventories maintained by different teams, reconciled by spreadsheet, with a measurable drift between the designed, built, configured and discovered state of the same asset.
  • Alarm systems per domain. Radio access, transport, IP, core and fixed access each with their own event formats, severities and suppression rules, and no shared normalisation.
  • Procedures that live in documents. Methods of procedure, runbooks and escalation trees written as PDFs and wiki pages, unversioned, occasionally contradictory, and never machine-readable.
  • Integration by exception. Point-to-point interfaces accumulated over fifteen years, several of them undocumented, some of them reading directly from a database that no longer has an owner.
  • A ticketing system as the only end-to-end truth. The one place where the whole incident lifecycle is visible, which is precisely why it is the best available substrate for process mining at the start of the journey.

This inheritance is not an argument for delay; it is an argument for sequence. The operator does not need a perfect inventory to start, but it does need to know how imperfect the inventory is, because that number determines which scenarios can be automated safely in year one. An agent reasoning over a topology that is 78% accurate will be confidently wrong about roughly one in five decisions, and confident wrongness executed at machine speed is the single most damaging outcome available in this programme.

Before assigning responsibilities, the target has to be specific, because different autonomy levels demand different organizations. The TM Forum scale runs from Level 0, where every reactive task depends on the judgement of the engineer on shift, to Level 5, where the system manages the full lifecycle of multiple services across multiple domains. The industry has converged on Level 4 as the meaningful commercial target, and the reason is that Level 4 is where the economics change rather than the tooling.

At Level 2 — partial autonomy — the system executes automations inside narrow network silos under strictly preconfigured engineering rules. The automation is real, the rules are human-authored, and the scope of each automation matches the scope of one team. Nothing about the operating model needs to change to reach Level 2, which is exactly why so many operators are there: it is the highest level attainable without renegotiating who decides what.

At Level 4 — high autonomy — the system makes predictive decisions driven by business intent, correlates events across physical silos and injects closed-loop corrective actions without immediate human intervention. Every clause in that sentence crosses an organizational boundary. Business intent has to be authored by someone with the authority to define what “good” means. Cross-silo correlation requires a shared data model that no single domain team owns. Closed-loop action without immediate human intervention requires a delegation of authority that, in most operators, has never been formally granted to anything other than a person.

The ETSI zero-touch framework provides the architectural vocabulary for that delegation, and it has matured quickly. The ISG ZSM body has published a study on the progression from automation to autonomy in ETSI GR ZSM 021 (May 2026), a study on the use of agents in autonomous networks in ETSI GR ZSM 020 (January 2026), a network-as-a-service framework in ETSI GR ZSM 019, and — most relevant to the governance question — a threat and risk analysis specific to closed loops in ETSI GR ZSM 017. Intent-driven autonomy has its own generic specification in ETSI GR ZSM 011. The existence of a dedicated closed-loop security study is itself the signal: the standards bodies treat autonomous action as a new attack surface, and so should the operating model.

The distance between ambition and attainment is wide enough to plan against. The measured and declared positions, separated from the projections, look like this:

  • Measured today. TM Forum’s 2026 benchmark of around 80 operators places roughly 4% of respondents at Level 4 overall, 17% at Level 3 and 31% at Level 2, with 21% at Level 3 or above overall against 19% the previous year (TM Forum, March 2026).
  • Declared intent. In the same benchmark, 20% expect to reach Level 4 or above by 2027 and 81% target it by 2030, with 75% increasing autonomous-network investment this year.
  • Self-assessment. 88% of respondents to the 2026 NVIDIA survey rate their own operation between Level 1 and Level 3, and 65% say AI is the primary force behind their network automation effort.
  • Modelled value. STL Partners models roughly US$794 million of annual capex, opex and revenue upside for an average operator moving up the autonomy curve — a projection, not a measured result (STL Partners).
  • Measured at Level 4. China Mobile’s Level 4 network operation centre case, published by TM Forum, reports an average 30% reduction in mean time to repair for faults and complaints, more than 30% reduction in backend operational effort, and an equivalent of 5,500 manual functions replaced (TM Forum case study).

Two readings matter for the operating model. First, the gap between 4% attained and 81% targeted by 2030 is a nine-in-ten failure rate against stated ambition unless something structural changes, and the surveys are consistent that the missing ingredient is not budget. Second, the one large public Level 4 case reports its gains in effort and repair time rather than in headcount, which is a reminder that the value of this programme is capacity released and outages avoided, and that the business case should be written in those terms from day one.

Unifying network and security operations is not an organizational fashion; it follows from the data. Both functions consume the same telemetry, both need the same topology and inventory truth to judge impact, both correlate events into incidents, and both increasingly act through the same orchestration layer. Building two separate agentic platforms over one network duplicates the expensive part — the context — and guarantees that the two operations reach different conclusions about the same event.

What must not be unified is the authority to act. A security containment action and a network optimisation action have different approval chains, different regulatory exposure and different reversibility. The workable design is a shared foundation with differentiated decision rights: one data layer, one API layer, one identity and logging regime, one scenario lifecycle, and separate risk tiers, approvers and escalation paths for security response. That distinction is the reason the decision-rights matrix in the next section has to be written down rather than assumed.

The model that works divides the work by the question each party is best placed to answer, not by the technology each party owns. Network Support answers what and why. The Integrator answers how it gets built and how fast. Systems answers where it runs and under what controls. Every recurring dispute in these programmes can be traced to one of those three questions being answered by the wrong party.

The three-party model. Network Support owns the what and the why, the Integrator owns the build and part of the result, Systems owns the how and the where.

The functional area that runs the operation becomes the product owner of the transformed NOC/SOC, and that is a genuine change of role rather than a title. Being a product owner means holding a prioritised backlog, saying no to things, defining what success means numerically, and personally accepting or rejecting the result. In practice the mandate covers five decisions:

  • Which scenarios enter the portfolio and in what order, based on volume, effort, customer impact and contractual exposure — not on which one is technically interesting.
  • The intent and the service quality indicators that each scenario must move: mean time to detect, mean time to repair, repeat-fault rate, truck rolls avoided, complaint volume for a given service class.
  • The tolerable risk envelope: which actions may ever be executed autonomously, on which network elements, in which maintenance windows, and with what maximum affected-customer count.
  • Acceptance of realised value, which means signing that the efficiency claimed by a deployed capability is the efficiency the operation actually experienced.
  • Retirement: which capabilities are corrected, demoted to a lower autonomy step, or switched off because they no longer earn their maintenance cost.

The most commonly skipped item on that list is the last one. Portfolios of automations grow monotonically because nobody is accountable for removing them, and the maintenance drag of obsolete automations is what eventually makes the operation slower than it was before. A product owner who has never retired anything is not exercising the role.

The partner’s contribution is a permanent capability to analyse, redesign, build, test and continuously improve, which is a different contract from delivering a specified list of automations. The distinction is observable in the first meeting of the week: a ticket-taking partner asks what to build next, a transformation partner arrives with process-mining evidence about where the operation is losing hours and proposes what should be built next. The scope that makes the second behaviour possible includes:

  • Continuous process analysis over real execution logs from the OSS and the ticketing system, rather than interviews about how the process is supposed to work.
  • Joint redesign with the operation before any automation is built, because automating an unfixed process industrialises the defect.
  • Construction and testing of agents, tools and connectors on the platform Systems provides, consuming approved APIs rather than inventing access paths.
  • Operation and tuning of the deployed loops, including the unglamorous work of prompt and retrieval maintenance, threshold recalibration and failure analysis.
  • Traceability between capability, operational change and euros, so that the efficiency claim can be audited rather than asserted.

What the Integrator must not hold is equally important. It does not select priorities, it does not define architecture unilaterally, it does not own the identities its agents use, and it does not decide when an automation is promoted to a higher autonomy step. A partner that accumulates those four things becomes structurally irreplaceable, which is bad for the operator’s cost base and bad for the partner’s incentive to industrialise.

The IT and systems area stops being an approver of individual initiatives and becomes the supplier of the environment in which initiatives are safe by construction. This is the hardest cultural shift of the three, because it converts a veto into a product, and products are judged by adoption. A Systems function whose platform is bypassed has failed, regardless of how rigorous its reviews were. The mandate is concrete:

  • Standards and catalogue for platform, identity, data, integration, observability and lifecycle, published as usable assets rather than policy documents.
  • Reusable patterns — approved connector templates, a vetted retrieval pattern, a standard action-and-rollback wrapper — precisely so that repetitive reviews become unnecessary.
  • Environments: the sandbox, the test data, and the digital twin or simulation capability where changes are validated before they touch production.
  • Runtime guardrails: identity issuance and rotation, policy enforcement, prompt-injection defence, immutable logging and cost telemetry.
  • Reinforced intervention — genuine deep review — reserved for core systems, write paths into corporate systems, new integrations and high-risk actions.

The measure of a healthy Systems function in this model is the ratio of pre-approved paths to case-by-case reviews. If nine out of ten scenarios can be built entirely on published patterns and only the tenth needs a conversation, the platform is doing its job. If the ratio is inverted, the operator has renamed the gate rather than redesigned it, and the fast track will exist only on the slide.

A practical diagnostic: take any decision in the programme and try to assign it in one sentence without a conjunction. “Network Support decides which scenarios are built.” “Systems decides which identity an agent receives.” “The Integrator decides how the agent is decomposed into tools.” If the sentence needs an “and” between two parties, the decision has not been assigned — it has been scheduled for a meeting. Programmes accumulate those unassigned decisions until the meeting load becomes the actual operating model, which is how a transformation designed to remove handoffs ends up creating new ones.

Role descriptions are agreed easily and interpreted differently, which is why the operating model has to be written as a matrix of specific decisions. The version below assigns four parties — Network Support, Systems and IT, Network Engineering as the domain authority, and the Integrator — across the decisions that actually cause friction. One party is accountable per row; several may be responsible for executing.

The decision-rights matrix. Exactly one accountable owner per row, and the rows that generate the most conflict are the ones about acting on the live network.

Each line below states the decision, the accountable owner and the reason the accountability sits there. The reasoning matters more than the letter, because it is what allows a new decision to be assigned by analogy later:

  • Business intents and scenario priorities — Network Support. The party accountable for service quality and cost is the only one that can rank the queue without being influenced by build effort.
  • AI platform architecture — Systems. Architecture decisions outlive scenarios and partners; whoever will still be operating the platform in five years must own its shape.
  • Agent construction and deployment — the Integrator. Accountability for the build must sit with the party that has the engineering capacity, or the build becomes a committee output.
  • Event correlation and alarm enrichment logic — Network Support. The definition of what constitutes an incident is an operational judgement, even when the correlation is implemented by data engineers.
  • Simulation and validation of changes in the twin — Network Engineering. The domain authority owns the fidelity of the model against which changes are tested, because it owns the consequences of the model being wrong.
  • Active remediation and command execution on the network — Network Engineering. The authority to change network state does not transfer to whoever wrote the automation.
  • Production go-live of an automation — Network Support. Go-live is an operational risk acceptance, and risk acceptance belongs to the party that carries the service.
  • Agent identity and credential governance — Systems. Machine identity is a security control, and security controls are never delegated to the party being controlled.
  • Audit of token, inference and infrastructure cost — Network Support, executed by Systems. The consumer of the capability has to see its unit cost, or the economics of autonomy are never questioned.
  • Root-cause analysis of major incidents — Network Support. Including incidents caused by an automation, which is the row most operators forget to write down.

The first is production go-live, which is routinely assigned to Systems by default. The instinct is understandable — Systems has the deployment tooling — but it produces the universal-approver pathology described earlier, and it separates the risk decision from the risk owner. The workable split is that Systems is responsible for the technical gate (the automation complies with the platform standards, has an identity, is logged, has a rollback) while Network Support is accountable for the operational gate (this behaviour is acceptable on my network, on these elements, at this hour). Two gates, two owners, one release.

The second is remediation authority, which is frequently left implicit and then discovered during an incident. When an agent executes a corrective action, the operator needs a documented answer to a simple question: on whose authority did that command run. The answer must be a named human role that authorised the class of action in advance, with the agent acting as a delegated executor inside a defined envelope. Any other answer — including “the platform allowed it” — is an audit finding waiting to happen and, in regulated markets, a compliance exposure.

A matrix without an escalation rule is decorative. The rule that works is simple and should be written into the governance charter: the accountable party decides, the responsible party may record a formal objection, and an unresolved objection about safety or architecture suspends the change until the governance committee meets. The asymmetry is deliberate — an objection can stop a change but cannot force one — because the failure mode being prevented is a fast operation overruling a slow control, not the reverse.

The second half of the rule is a time limit. A suspension that lasts until the next scheduled committee is a de facto veto if the committee meets quarterly, which is one of the reasons the governance rhythm proposed later is fortnightly with an emergency path. Control that is proportional to risk also has to be proportional in speed, or the organization will simply route around it.

Once decision rights are clear, the operator has to choose a structure to execute them, and there are only three serious candidates. Each has a coherent logic and a characteristic failure. The choice is not a matter of taste: it determines how many scenarios the operation can run concurrently and how quickly a pattern built in one domain reaches the others.

The recommended shape. A platform team owned by Systems supplies APIs, identities, sandbox and observability; stable squads aligned to network value streams do the delivery with a co-managed backlog.

A single central team owns the models, the platform and the delivery of every use case across all domains. It is the fastest structure to stand up, it concentrates scarce data-science talent, and it produces consistent engineering. It also has a hard ceiling: the central team becomes the sole path to production, and its throughput caps the programme regardless of how much demand the domains generate.

The deeper problem is distance from the operation. Agentic automation depends on tacit operational knowledge — which alarm is habitually noise on this vendor’s equipment, why this procedure has a manual step, which customer circuits nobody touches on a Friday. A central team acquires that knowledge through interviews and loses it between use cases, which is why centre-of-excellence portfolios tend to be rich in demonstrations and thin in production loops. The archetype is a reasonable starting point for the first six months and a poor destination.

Each domain — radio access, transport, IP, core, fixed access, security — builds its own automation with its own tooling and its own partner. Proximity to the operation is excellent and delivery is fast in the first year. Then the structural cost appears: five ways of representing an alarm, five identity schemes, five retrieval patterns, and no cross-domain correlation, which is precisely the capability that separates Level 2 from Level 4.

Full federation recreates the silos the transformation was meant to dissolve, one automation platform at a time. It also multiplies the security surface: every domain negotiates its own access paths to network elements, and the operator loses the single audit trail that a regulator or an internal auditor will eventually ask for. The archetype is attractive because it requires no central agreement, which is exactly why it fails at the point where central agreement becomes necessary.

The hybrid separates what must be common from what must be close to the operation, and it is the only one of the three that scales without either bottleneck or fragmentation. Systems runs a platform team that publishes self-service capabilities — the API catalogue, the knowledge graph, agent identity issuance, the sandbox and twin, the observability and cost telemetry. Delivery happens in stable multidisciplinary squads aligned to operational value streams rather than to technology towers, in which engineers from Network Support and developers from the Integrator work on one shared backlog. The pattern maps directly onto the platform-team and stream-aligned-team distinction popularised in Team Topologies, and the mapping is useful because the failure modes documented there are the ones operators hit.

Three properties make this structure work, and all three are frequently dropped in implementation. Squads must be stable — the same people across scenarios, so tacit knowledge accumulates. The backlog must be genuinely shared, in one tool, with one ranking, not two synchronised plans. And the platform must be consumable without a ticket to the platform team, because a self-service platform that requires a request is just a gate with better branding.

A squad that can take a scenario from discovery to closed loop needs six roles, and it is smaller than most operators expect. Six to nine people is the working range; beyond that, coordination cost eats the proximity advantage:

  • An operations engineer from Network Support who has actually worked the shift and carries the tacit knowledge — the scarcest and most important seat.
  • Two to three engineers from the Integrator covering agent assembly, tool and connector development, and test automation.
  • A data engineer accountable for the data products the scenario consumes and for closing the gaps it exposes in inventory and topology.
  • A domain specialist from Network Engineering, part-time, who validates that a proposed action is technically sound and owns the simulation fidelity.
  • A platform contact from Systems, part-time, embedded enough to unblock API, identity and sandbox questions in hours rather than sprints.
  • A squad lead pairing — one from Network Support, one from the Integrator — because single leadership in a two-organization team always resolves into one side directing the other.

The composition encodes the decision-rights model in seating. Every squad contains someone who can say what the operation needs, someone who can build it, and someone who can say whether the platform allows it — which is why the vast majority of decisions never need to leave the squad. Governance forums then handle the residue rather than the flow, and the residue is where scarce senior attention belongs.

Between the head of operations and the squad there is a missing job, and its absence is the most common structural defect in these programmes. Somebody has to convert operational pain into ranked, specified, measurable scenarios; maintain the value case for each one; and hold the line when a stakeholder wants their pet automation prioritised. That is a product management function inside Network Support, staffed with people who understand operations rather than with programme managers who track milestones.

When the role is unfilled, the Integrator fills it by default, and the model quietly inverts. The partner starts proposing the priorities, the operation starts reviewing proposals rather than setting direction, and within two quarters the operator has outsourced the one thing it cannot outsource: the definition of what its operation should become. Two or three named people with this responsibility, funded from the operations budget, are worth more to the outcome than a doubling of the build team.

The three-party model becomes operational when it is projected onto the technical layers, because the same party plays a different role at each one. Three layers are enough to describe the whole operation: data and inventory at the bottom, enablers and intelligence in the middle, processes and use cases at the top. The value is visible only at the top, and the failures almost always originate at the bottom.

The layered demarcation. Each layer has one owner per column and no layer can be skipped: the data foundation determines what the process layer is allowed to automate.

This layer exists to eliminate the fragmentation and semantic inconsistency that historically kill automation programmes, and it has three components. A unified real-time inventory that reconciles physical, logical, service and customer records and tracks the lifecycle states of every asset — designed, built, configured, discovered. A network knowledge graph that maps the dependencies between physical elements, logical services and customer contracts, so an agent can answer “who is affected” without guessing. And a unified observability fabric that ingests and normalises telemetry, alarms, performance counters and traces under common standards such as OpenTelemetry.

The demarcation at this layer is unambiguous and rarely respected. Network Support owns the operational data: it is accountable for the fidelity of the inventory, because it is the party whose decisions degrade when the inventory is wrong. The Integrator consumes and reconciles: it builds the models and flows that expose inconsistencies and it repairs the topology systematically rather than case by case. Systems provides the machinery: the graph engine, the data platform, the ingestion pipelines and the governance of the data products themselves, following the “data as a product” discipline that data mesh formalised.

One rule at this layer prevents most of the downstream pain: no scenario is approved above the accuracy its data supports. If the topology for the transport domain is validated at 96% and the access domain at 74%, transport scenarios can move toward bounded autonomy while access scenarios stay at recommendation until the data is repaired. Making data quality a gating parameter rather than a background complaint is what turns inventory remediation from a thankless chore into a funded prerequisite with a visible payoff.

The middle layer is the reasoning engine and the secure communications bus, and it is where the operator’s leverage is highest because everything above it is reusable only if this layer is standard. It has three components: an integration platform based on open APIs, a multi-domain automation orchestrator that abstracts vendor-specific controllers and consoles behind a normalised execution interface, and the guardrail machinery — identity, policy and prompt-injection defence — that makes execution on a live network acceptable.

Standardisation at this layer is now a realistic choice rather than an aspiration. TM Forum’s Open API programme provides the intent interface in TMF921, and the agentic work — an agent interface for declaration, discovery and coordination, and an assistant interface for natural-language interaction — is being specified and tested through the Forum’s Agent Fabric catalyst work as TMF939 and TMF785. These are emerging rather than settled specifications, and the operator should treat them as a direction of travel to align with rather than a completed standard to procure against. The tool-invocation layer beneath them is converging on the Model Context Protocol for exposing systems to agents in a uniform way.

The demarcation here is the one that most often breaks under delivery pressure. Systems administers the API catalogue, issues and rotates credentials and provides the sandbox. The Integrator assembles agents, maps their calls to the approved APIs and builds the connectors. Network Support defines the access policy by role and the risk criteria the guardrails enforce. The rule that keeps it honest is absolute: no direct access to network element consoles or to underlying databases, ever, regardless of deadline. Every exception granted at this layer becomes a permanent integration that nobody will remove and that the next audit will find.

The top layer is the detect-to-resolve loop itself, and it decomposes into three stages that map cleanly onto agent responsibilities. Sense: the observability analytics monitor service-quality indicators, translate raw telemetry into customer-experience impact, suppress repetitive noise through automatic correlation and surface silent incidents that never raise a physical alarm. Decide: a diagnostic agent queries the knowledge graph, correlates dependencies, applies digitised procedures and determines the actual fault rather than the most visible symptom. Act: a remediation agent designs a restoration strategy from business intent, the orchestrator executes the change, and post-validation runs continuously against the metrics.

The demarcation at the top layer is where the operating model earns its keep. Network Support ranks the scenarios and defines the customer-experience indicators that each loop must move. The Integrator tests, deploys and continuously tunes the loops in production. Systems applies the runtime guardrails, audits the traces and monitors the infrastructure and inference consumption. Nobody’s role changes between scenarios, which is what allows the second scenario to be delivered faster than the first — the whole point of an operating model as opposed to a project.

No layer may be skipped, and the temptation to skip is strongest exactly where skipping is most expensive. A digital twin built over an inaccurate inventory produces confident simulations of a network that does not exist. An agent orchestration layer built over point-to-point integrations inherits every fragility of those integrations and adds non-determinism on top. The order is not a preference: process and data before automation, automation before orchestration, orchestration before agents, agents before closed loops.

The practical enforcement mechanism is a foundation budget that is not raidable. Data reconciliation, ontology work and API publication have no demo value and are the first line cut when a steering committee wants a visible result. Ring-fencing that budget — and reporting graph accuracy alongside scenario counts in the same monthly pack — is what stops the programme from spending year one building demonstrations on top of a foundation that cannot support year two.

The single most consequential change in the operating model is what the organization counts as a piece of work. Traditional demand management counts requests: isolated automation scripts, reports, integrations, each with its own justification and its own end date. The target model counts end-to-end operational scenarios — “restore service after a mass access-network degradation”, “diagnose and clear repeat faults on a fibre segment”, “contain a lateral-movement alert in the management network” — each with a named owner, a measured baseline and a permanent lifecycle.

The six-phase lifecycle. Every scenario runs the same loop, with an explicit owner per phase and kill criteria that allow the portfolio to shrink as well as grow.

Candidate scenarios are found by mining the operation’s own execution records, not by asking teams what hurts. Process mining over OSS transaction logs and ITSM ticket histories reconstructs how work actually flows: where cases wait, which steps repeat, how often an incident bounces between two teams, and which activities consume the most engineer-hours per resolved case. The output is a ranked inventory of bottlenecks with volumes and durations attached, which is a far better starting point than a workshop, because interviews reliably surface the most annoying process rather than the most expensive one.

Ranking uses four factors, and the fourth is the one operators forget. Volume of occurrences, human effort per occurrence, customer or SLA exposure, and data readiness — the accuracy of the inventory, topology and telemetry the scenario would depend on. A high-value scenario sitting on 70%-accurate data belongs in the queue behind the data repair that unblocks it, and saying so explicitly is what converts data quality from an excuse into a scheduled task with a beneficiary.

Automating a broken process produces a faster broken process, and this phase is where the joint work between Network Support and the Integrator is most valuable. The redesign removes steps that exist only because a system could not talk to another system, collapses approval hops that no longer carry information, and normalises how priority is derived so that it reflects real customer impact rather than the severity field a vendor’s element manager happens to emit.

The technical output of the redesign is a semantic layer over the legacy estate. An abstraction that unifies context across inherited systems and gives the agents a stable vocabulary — service, site, circuit, customer, severity, impact — decoupled from the fifteen ways those concepts are represented underneath. This is unglamorous ontology work and it is the difference between an agent that generalises across domains and one that has to be rebuilt for every system it meets.

The Integrator’s factory translates the redesigned process into an orchestration of specialised agents and tools, and the emphasis belongs on assembly. A scenario is composed from components that already exist — an impact-prioritisation agent, a telemetry query agent, a remediation agent bound to a set of digitised methods of procedure, a retrieval layer over the knowledge base — configured for this process rather than written for it. The measure of a maturing factory is the ratio of reused components to new ones, and it should rise every quarter.

Validation happens in the sandbox and, for anything that touches network state, in simulation before production. This is where the operator discovers whether its digital twin is a decision-grade asset or a diagram: a twin that cannot reproduce the failure mode being automated is not a validation environment. Scenarios whose actions cannot be simulated credibly are capped at the recommendation step of the autonomy ladder until they can be, which is a slower path but the only defensible one.

A scenario enters production at the lowest useful autonomy step and climbs only on evidence. In the first weeks the agent proposes and a human executes; the comparison between proposed and executed actions is the training data for the trust decision. Every action that touches configuration is wrapped in the standard control envelope: a configuration snapshot before injection, continuous post-validation of the local metrics, and automatic rollback if degradation appears within the defined window.

Running is a shared responsibility and should be staffed as one. The Integrator tunes the loop; Systems watches identity, policy and cost telemetry; Network Support watches whether the operation is actually getting better and retains the authority to demote the scenario a step at any time without a committee. That last provision matters more than it appears: an operation that cannot unilaterally reduce autonomy will refuse to grant it in the first place.

Every scenario carries a measurement obligation that includes its own cost, which is what separates this model from a conventional automation programme. Operational effect is tracked against the pre-automation baseline — detection time, repair time, human touches per incident, repeat faults, field visits avoided, engineering hours released. Economic cost is tracked per transaction: inference and token consumption, orchestration and infrastructure, and the human time still required to supervise.

Cost per outcome is the metric that keeps the portfolio honest. A diagnostic loop that resolves a case for less than the loaded cost of the equivalent manual work is a business; one that costs more and is retained because it is strategic is a subsidy, and the operator should at least know which of the two it is running. Making this visible per scenario, monthly, is also the most effective defence against the cost-driven cancellations that Gartner identifies as a leading cause of agentic programme failure.

Automation failures and human overrides are the highest-value data the operation generates, and most operators discard them. Every case where an engineer corrected an agent, every rollback, every escalation the loop failed to resolve should be analysed on a fixed cadence, with the conclusion written back into the knowledge base, the procedures and the guardrails. The improvement work needs a named owner and dedicated capacity in the Integrator’s team, because unfunded improvement is always deferred in favour of the next scenario.

The lifecycle also has to allow scenarios to leave. Kill criteria stated at the start — no measured baseline, reliability still short of target after two improvement cycles, cost per outcome persistently above the manual alternative — give the portfolio a mechanism to shrink. A portfolio that only grows is not a portfolio; it is an inventory of technical debt with a governance forum attached.

The governance change that unlocks throughput is replacing one uniform gate with a small number of pre-approved routes, chosen by what the automation can actually do. The classification has to be mechanical rather than negotiated: an engineer should be able to determine a scenario’s tier from its own properties in minutes, without a meeting, and the tier determines the route, the evidence required and the approver.

Control proportional to risk. The tier determines the route and the approver; the autonomy ladder determines how much authority a proven scenario is allowed to accumulate.

Four tiers cover the realistic portfolio, and the boundaries are drawn by what the automation writes to and how far its effects propagate:

  • Tier 1 — read-only insight on the approved platform. Summarisation, correlation, natural-language querying of telemetry and topology, draft root-cause narratives. Route: the squad decides, registration in the catalogue is mandatory, no external approval. Target lead time from idea to production: days.
  • Tier 2 — writes confined to operational tooling. Creating, enriching, correlating and closing tickets, updating workflow state, scheduling field work. Route: fast track with a scheduled post-implementation review, because the blast radius is process disruption rather than network damage.
  • Tier 3 — configuration change inside one bounded domain. Parameter adjustment, service restart, traffic rebalancing, port or channel reconfiguration within a defined envelope. Route: reinforced Systems and Network Engineering review, simulation evidence required, snapshot and rollback mandatory.
  • Tier 4 — core, security policy or multi-domain action. Anything touching the core, changing security posture, crossing domain boundaries in a single action, or affecting more than a defined customer threshold. Route: governance committee decision, staged rollout, explicit human authorisation per action class.

The distribution of a healthy portfolio is heavily weighted to the first two tiers, and that is the point. If most scenarios are Tier 1 and Tier 2, most of the portfolio moves at squad speed, and the reinforced review capacity — which is a genuinely scarce resource made of senior architects and security engineers — is spent entirely on the Tier 3 and Tier 4 minority where it changes outcomes.

The disqualifiers should be short, absolute and known in advance, so that nobody designs a scenario into a slow lane by accident. Five conditions move a scenario out of the fast track regardless of its apparent simplicity:

  • It writes to a core system — network core, billing, customer records, security policy stores.
  • It introduces a new integration not already published in the API catalogue, because a new access path is an architectural decision.
  • It processes personal or regulated data beyond what the approved data products already expose.
  • It acts autonomously at any step above assisted execution, which is a promotion decision rather than a build decision.
  • Its user population widens materially — a tool built for one shift team being published to the whole operation is a different risk object with the same code.

The last disqualifier is the one that most governance frameworks miss. Scope creep in an agentic tool is silent: the same automation, unchanged, becomes materially riskier when its audience grows tenfold, because the assumptions the original users held in their heads stop being universal. Treating publication scope as a governed attribute — with a re-review when it changes — closes a gap that code review alone will never catch.

Autonomy is not a property of a scenario; it is a level of authority granted to a scenario after it has demonstrated reliability, and it is granted in four steps. At recommend, the agent proposes and a human executes. At assisted, the agent prepares and executes with per-action human approval. At bounded, the agent acts autonomously inside an explicitly defined envelope — these element types, this parameter range, these hours, this maximum affected-customer count — and escalates anything outside it. At closed loop, the agent acts and humans audit after the fact.

Promotion between steps requires evidence and a decision, never elapsed time. The evidence pack for a promotion is specific: the agreement rate between proposed and human-executed actions over a meaningful sample, the false-positive rate of the underlying diagnosis, the rollback success rate in simulation and in production, the incident record of the scenario to date, and the cost per outcome. The governance committee decides; the operation can demote unilaterally. That asymmetry — hard to promote, easy to demote — is what makes it rational for an operations leader to allow the first autonomous action at all.

Reinforced review is only credible if it is fast enough to be used rather than avoided. The commitment should be explicit and published: Tier 3 review completed within a defined number of working days, Tier 4 heard at the next fortnightly committee with an emergency path for incident-driven cases. A control function that does not publish its own turnaround times is asking the organization to accept an unbounded delay, and organizations respond to unbounded delays by building outside the process.

The principle underneath all of this deserves stating plainly, because it is routinely misread in both directions. A fast track is not the absence of governance; it is governance decided in advance, encoded in the platform, and applied automatically. The pre-approved path carries mandatory registration, an owner, a machine identity, full logging and a defined retirement route. What the fast track removes is the queue, not the control.

If Systems is going to stop reviewing every initiative, it has to deliver a set of controls that make the fast track defensible, and six of them are non-negotiable for a network that carries live traffic. These are the deliverables that justify the platform team’s existence, and each one should be consumable as a service rather than described in a policy.

Every agent interaction with the network and with corporate systems executes through the standardised open API layer, with direct access to element consoles and databases prohibited without exception. The prohibition is what makes everything else enforceable: identity, rate limiting, policy, logging and cost attribution are all applied at the API boundary, and an automation that bypasses the boundary is invisible to all of them. The corollary is an obligation on Systems — if the catalogue lacks the API a scenario needs, publishing it becomes a priority rather than a reason to grant an exception.

Agents reason over the corporate network graph rather than over free-text documents, and that architectural choice is a safety control, not a performance optimisation. A model asked to infer topology from documentation will produce plausible relationships that do not exist; a model that queries a graph unifying physical and logical topology in real time either finds the relationship or does not. The graph is also what makes impact analysis deterministic — which customers, which services, which contracts — and impact determines whether an action is Tier 3 or Tier 4.

Every agent holds its own cryptographic machine identity, scoped to least privilege and rotated per task, replacing the generic automation accounts that most operators still use. This is the control that makes autonomous action auditable at all: without per-agent identity, the log shows that “the automation platform” changed a parameter, which is useless in a post-incident review and unacceptable in a security investigation. It is also the control that makes revocation possible — a misbehaving agent can be disabled without disabling every automation sharing its credentials.

Network telemetry is untrusted input, and in an agentic operation it becomes an attack surface that did not previously exist. Syslog messages, SNMP traps, device banners, ticket free-text and vendor alarm descriptions all flow into prompts, and any of them can be crafted to manipulate an agent’s reasoning. A stateful inspection layer between raw network variables and the templates that consume them — sanitising, delimiting and validating before injection — is the control ETSI’s closed-loop threat analysis exists to formalise. Treating this as an academic risk is a choice the operator only gets to make once.

Every agent action is recorded under open telemetry standards with a unique transaction identifier that ties together the trigger, the reasoning trace, the tools invoked, the parameters used, the token cost and the outcome. The requirement is driven by three separate consumers: incident forensics needs the causal chain, finance needs the cost attribution, and audit needs immutability. Building the logging to satisfy only the first produces a system that fails the other two at exactly the moment they are needed.

No change reaches a physical network element without a prior configuration snapshot and an armed rollback, and the rollback triggers on metric degradation rather than on human noticing. Alongside it sits an explicit blast-radius limit per action class — how many elements, how many customers, how much of a region — enforced by the platform rather than by the agent’s own judgement. The pairing matters: the snapshot bounds the recovery time, the blast-radius limit bounds the damage, and neither substitutes for the other.

One further control belongs in the same category and is easy to forget until it is needed: the kill switch. A single, tested, documented mechanism that suspends all autonomous action across the estate, exercisable by the duty manager without escalation. It will be used rarely, possibly never, and its existence is what allows an operations director to sign off on closed-loop execution in the first place. Testing it on a schedule, like any other emergency procedure, is what keeps it real.

An operating model is real only in its calendar, and three recurring events are sufficient to run this one. More forums than that and the model becomes a meeting schedule; fewer, and decisions accumulate until they are made under incident pressure. Each event has one purpose, one decision type and a fixed attendance.

One board, one ranking, visible to all three parties, containing every scenario candidate and every foundation task. Network Support ranks by operational and commercial impact, the Integrator attaches feasibility and effort, Systems attaches the integration and control requirements each item carries. The single most important property is that data-foundation work and scenario work compete in the same list, because a separate “technical debt backlog” is a backlog that never gets funded.

Two disciplines keep the board useful. Every item carries its measured baseline before it is eligible to be started — no baseline, no build — and every item carries a named accountable owner from Network Support, not a team name. Boards without both properties degrade within a quarter into a list of aspirations with no way to tell which ones are progressing.

A short morning session focused exclusively on how the closed loops behaved in the last twenty-four hours, not on project status. The agenda is fixed: which loops executed, which failed, which triggered rollback, which were overridden by an engineer, and what the token and infrastructure consumption looked like. Anything requiring more than a two-minute discussion is taken offline with a named owner and a date.

The purpose of the ritual is to prevent the most common regression in autonomous operations: silent deactivation. When an automation misbehaves and nobody owns the fix within a day, the shift team disables it and reverts to manual work, and the capability is lost without any decision ever being recorded. A daily forum that surfaces failures and assigns a correction immediately — usually a fix in the retrieval content, the procedure or the threshold rather than in the model — is what keeps the autonomy the operator has already paid for.

The forum where the three areas review the evolution of autonomy and make the decisions that squads cannot make alone. Its decision set is narrow and should stay narrow: authorising promotions on the autonomy ladder, approving Tier 4 scenarios, clearing exceptions to the platform standards, unblocking cross-area dependencies and releasing the next wave of work. It does not review individual builds, and it does not approve fast-track scenarios — if it does, the fast track has been abolished by procedure.

The evidence pack is standard and short. Reliability and agreement statistics for the scenarios proposed for promotion, the rollback and incident record since the last meeting, the twin-validation results, the token and infrastructure consumption trend, and the value realised against baseline for scenarios in production. A committee that receives narrative slides instead of these five artefacts will approve on confidence rather than evidence, which is the specific failure the fortnightly cadence exists to prevent.

Once a quarter the question changes from “is it working” to “is it worth it”, and the audience widens to include finance. The review compares committed efficiency against realised efficiency scenario by scenario, examines cost per outcome trends, decides which capabilities are scaled, corrected or retired, and adjusts the next quarter’s investment. It is also where the transfer of capability from the Integrator to the operator is measured, which is the item most likely to be quietly dropped from the agenda if nobody insists on it.

Involving finance early is not a bureaucratic concession; it is a defence mechanism. The programmes that survive their second year are the ones where the finance function has seen the measurement method, agreed it, and can therefore defend the numbers when a cost-reduction cycle arrives. Efficiency claims that appear for the first time in a budget negotiation are treated, correctly, as unverified.

The contract with the Integrator determines its behaviour far more reliably than the operating model diagram does, and most of these programmes are commercially structured to produce exactly the behaviour they claim to be replacing. If the partner is paid for delivered effort against specified requests, it will optimise for delivered effort against specified requests, whatever the governance charter says about co-creation.

Effort-based contracting makes reuse commercially irrational for the supplier and makes scope discipline the operator’s sole responsibility. Under time and materials, a component built once and reused ten times generates one tenth of the revenue of ten bespoke builds, and the partner has no funded route to propose work that nobody asked for — which is precisely the behaviour the transformation depends on. The model is not dishonest; it is simply aimed at a different outcome.

The workable structure is a stable capacity core plus an outcome-linked component, in roughly that order of magnitude. A committed multidisciplinary team funded as capacity, which buys continuity and tacit knowledge, and a variable element tied to measured operational results — repair time, human touches, cases resolved without escalation, engineering hours released. The proportion tied to outcomes matters less than the fact that it exists and is measured by an agreed method.

Every outcome-linked contract in this domain fails at the same point: nobody measured the “before” precisely enough to make the “after” arguable. The baseline has to be established before any automation is deployed, agreed in writing by both parties and by finance, and defined at a granularity that survives scrutiny — which incident classes, which domains, which hours, which exclusions for major events. Establishing it takes weeks and is the highest-return administrative work in the entire programme.

The second requirement is an attribution rule agreed in advance. Repair times improve for many reasons — a network upgrade, a seasonal traffic change, a vendor firmware fix — and without a rule, every quarterly review becomes an argument about causation. Practical operators use domain-level control groups where the topology allows, staged rollouts that leave a comparable population unautomated for a defined period, or a fixed attribution split negotiated up front. Any of the three beats discovering the problem in month nine.

The intellectual property question in an agentic programme is not about models; it is about the operational knowledge that gets encoded during the work. The digitised methods of procedure, the ontology and semantic layer, the agent definitions and their tool bindings, the retrieval corpus, the evaluation sets and the test cases are the accumulated description of how the operator’s network is actually run. They must belong to the operator, be stored in the operator’s repositories, and be portable to another supplier.

The distinction to negotiate is between the partner’s generic accelerators and the operator’s specific knowledge. A partner’s reusable framework can reasonably remain the partner’s; the encoded procedures of this operator’s transport network cannot. Getting that boundary written into the contract before the first sprint is straightforward; getting it written after two years of accumulated artefacts is a commercial negotiation with a very weak hand.

Inference cost is the first genuinely variable operational cost most network operations functions have ever carried, and it behaves differently from everything around it. It scales with incident volume, it spikes exactly when the network is in trouble, it is sensitive to design choices invisible to the business — context window size, retrieval breadth, retry policy, how many agents review each other’s work — and it accrues in a cloud account that nobody in operations reads.

Three controls make it manageable and all three belong in the operating model rather than in a later cost-optimisation exercise. Per-scenario cost attribution through the transaction identifier, so consumption is visible where the value is claimed. A budget envelope per scenario with alerting, so a runaway loop is detected in hours rather than at month end. And an explicit efficiency-versus-accuracy review in the improve phase, because the cheapest reliable configuration is rarely the first one that worked.

A partnership deep enough to be valuable is deep enough to be dangerous, and the mitigation is designed reversibility rather than reduced commitment. Three provisions cover it: the artefacts are the operator’s and are continuously in the operator’s repositories; the platform, identities and API layer are operated by Systems, so the partner never controls the runtime; and a defined proportion of squad seats is held by operator staff who are rotated deliberately for capability transfer.

Capability transfer only happens when it is measured. A simple quarterly indicator — the share of scenarios in which the operator’s own engineers can perform the change without the partner — turns an aspiration into a tracked outcome. Left unmeasured, it does not occur, and the operator discovers at renewal that its operational knowledge lives in someone else’s team.

The measurement system has to do two jobs at once: prove that the operation is improving, and make each party’s contribution visible enough to be managed. A single blended KPI does neither. The workable set splits into operational, platform and economic indicators, each with an owner and a stated baseline.

These are the metrics the business already understands, measured against a baseline captured before automation:

  • Mean time to detect and mean time to repair, segmented by incident class and domain, since blended averages hide exactly the cases that matter.
  • Human touches per incident, the cleanest available proxy for how much of the process is genuinely automated end to end.
  • End-to-end automation rate: the share of incidents of a given class resolved with no human intervention, reported per class rather than as one number.
  • Repeat-fault rate and field visits avoided, which capture the predictive value of the operation rather than its reactive speed.
  • Engineering hours released, tracked as capacity redeployed to design and improvement work rather than as headcount removed.

These are the metrics that determine whether the fast track deserves to keep existing:

  • Topology and inventory accuracy, measured per domain against physical audit samples, and published as the gating parameter for autonomy promotion.
  • Rollback success rate, in simulation and in production — the number that determines how much autonomy the organization can rationally grant.
  • Share of agent traffic through approved APIs, where anything below 100% is an open finding rather than a metric.
  • Blocked-injection rate and prompt-firewall coverage across the input paths that feed agent context.
  • Time from scenario request to available platform capability, which measures the platform team as a supplier rather than as a gate.

Two numbers carry the business case and both must be reported per scenario. Cost per outcome — the fully loaded cost of resolving one case through the automated path, including inference, infrastructure and residual human supervision, compared with the manual baseline. And realised versus committed efficiency, which is the only number that determines whether the next wave gets funded.

Targets should be set from public reference points but never inherited from them. The Level 4 case published by TM Forum reports around 30% mean-time-to-repair improvement and more than 30% backend effort reduction, and those are useful order-of-magnitude anchors. They are not transferable: a different starting point, scope, technology mix, incident profile and metric definition will produce a different number, and any business case built by importing another operator’s percentage will be wrong in a direction nobody can predict.

Three legacy indicators actively damage this operating model and should be removed from the reporting pack rather than supplemented. Tickets closed per engineer rewards volume in a system whose goal is fewer tickets. Automations delivered rewards output over outcome and is the metric that produces the pilot graveyard. Project milestones met measures a programme’s adherence to a plan that was written before anyone knew what the data would support. Keeping them alongside the new set guarantees that the organization optimises for the familiar ones.

The sequence below assumes the realistic starting position described earlier: multiple inventories, per-domain alarm systems, procedures in documents and a decade of point-to-point integrations. Each quarter has an exit gate, and the gate is a condition rather than a date — a programme that passes a gate on schedule without meeting the condition has simply moved the failure later.

Twelve months with gates rather than dates. Each quarter has an exit condition, and the programme does not proceed on schedule alone.

The first quarter produces agreements and numbers, not automations, and resisting the pressure to demonstrate something is the main leadership task. The decision-rights matrix is written and signed by the three areas. The risk tiers and the autonomy ladder are defined with their approvers. Systems publishes the first version of the API catalogue and stands up the sandbox. Network Support staffs the operations product management role and captures the baseline: detection and repair times by incident class, human touches per incident, current automation rate, and a measured accuracy figure for inventory and topology per domain.

Process mining runs in parallel and produces the first ranked scenario list. The gate at the end of the quarter is blunt: no baseline, no build. An organization that cannot state its current mean time to repair by incident class has no way to prove any of the value it is about to claim, and will spend year two arguing about it.

The second quarter builds the data layer while delivering the first visible capability, and the pairing is deliberate. Inventory and topology reconciliation starts in the domain with the best data, the first network data products are built for inventory, alarms and performance, and the knowledge graph is populated for one or two domains rather than all of them. Simultaneously, two Tier 1 scenarios go live — typically alarm correlation with impact summarisation, and natural-language querying of topology and telemetry — because a foundation quarter with no user-visible output loses its sponsor.

The gate is a measured graph accuracy threshold in the target domains. Setting it explicitly — above 95% correspondence with physical reality on the entities the scenarios depend on — converts the perennial “our data is bad” complaint into a tracked engineering objective with a completion criterion, and it defines which domains are allowed to progress toward action in the following quarter.

The third quarter is where the operating model is tested, because it is where software starts proposing changes to a live network. Five or six scenarios run with a human in the loop, accumulating the agreement statistics that a promotion decision needs. The snapshot-and-rollback envelope is exercised repeatedly in the twin and then in production on low-risk changes. The prompt-injection defence, machine identities and per-action logging move from designed to verified.

The first bounded autonomy is authorised at the end of the quarter, for one scenario, in one domain, inside a narrow envelope. The gate is that rollback works every single time it is invoked in testing — not most of the time. An operator that grants autonomy before that condition is met is relying on the automation being right, which is precisely the assumption the entire control framework exists to avoid.

The fourth quarter tests whether anything built so far is reusable, which is the only question that determines whether year two costs less than year one. Patterns proven in the first domain are replicated to a second and a third, and the honest measure of the platform is how much of each new scenario is configuration rather than construction. The cost-per-outcome measurement is published per scenario, and the first efficiency acceptance is signed by Network Support.

The gate is value accepted, not capability delivered. A programme that reaches month twelve with capabilities in production and no signed acceptance of realised value has not proven anything to the people who will decide its second-year budget, and in the current cost environment that is how a technically successful programme gets cancelled.

The failure modes are consistent enough across operators to be treated as a checklist, and each has a countermeasure that belongs in the design rather than in the recovery plan:

  • Agents on top of broken data. A probabilistic decision layer over fragmented context produces fast, confident, wrong actions. Countermeasure: data accuracy as the gating parameter for autonomy, measured per domain.
  • The integrator drifts into priority setting. Happens whenever the operations product management role is unstaffed. Countermeasure: fund the role from the operations budget before the first sprint.
  • Systems keeps the veto and renames it. The fast track exists on the slide but every scenario still passes through one queue. Countermeasure: publish the ratio of pre-approved to reviewed scenarios monthly.
  • Autonomy granted by enthusiasm. A successful demonstration becomes a closed loop without an evidence pack. Countermeasure: promotion requires the five artefacts, and the committee is the only body that can grant it.
  • Silent deactivation. Shift teams disable misbehaving automations and nobody records it, so the programme reports capabilities it no longer has. Countermeasure: the daily run sync, and a live catalogue that shows the actual state of every loop.
  • Benefits promised without a baseline. Reference percentages from other operators imported into a business case. Countermeasure: no build without a measured baseline, and an attribution rule agreed in writing.
  • Inference cost discovered at year end. The variable cost of autonomy accrues invisibly until it becomes a cancellation argument. Countermeasure: per-scenario cost attribution and budget envelopes from the first scenario.
  • The organization outruns its people. Engineers who distrust the outputs re-verify everything manually and erase the gains; managers measured on headcount resist role conversion. Countermeasure: shadow-mode trust building, explainability as a build requirement, and retraining committed publicly and early.

The last one is the least technical and the most predictive. Every element of this operating model asks someone to give up something they currently control — the operation gives up doing the work by hand, Systems gives up the universal veto, the partner gives up open-ended scope. Programmes that treat those trades as change-management overhead rather than as the substance of the design are the ones that produce excellent architecture diagrams and a NOC that still runs on people reading alarm lists.

The model that lets a telco build an agentic NOC/SOC without either paralysis or recklessness is a hybrid, layered target operating model with segmented governance, and it can be stated in one paragraph. Network Support owns the process and the value: it ranks the scenarios, defines the intents and risk limits, and accepts or rejects the efficiency achieved. The Integrator is the transformation engine: it analyses, redesigns, builds, runs and improves inside the agreed guardrails, and carries a share of the efficiency commitment. Systems is the guarantor of the platform: it supplies the APIs, the graph, the identities, the sandbox and the observability, and reserves its deep review for the core, the write paths and the high-risk actions. Delivery happens in stable squads over a common platform, work is organized as a continuous portfolio of operational scenarios rather than a queue of requests, and autonomy is granted step by step against evidence.

The order of construction is not negotiable, and it is where most of the recoverable value is won or lost. Decision rights before platforms; measured baselines before builds; data and topology before automation; automation before orchestration; orchestration before agents; agents before closed loops. Every operator that has tried to compress that order has arrived at the same place — a portfolio of demonstrations that cannot be promoted because the foundation underneath them does not support the evidence that promotion requires.

The economics justify the discipline. With roughly 4% of operators at Level 4 today, 81% targeting it by 2030 and more than 40% of agentic projects across industries forecast to be cancelled before 2028, the differentiator is not access to models — every operator has that — but the ability to convert a pilot into a governed, measured, reusable capability. The one large published Level 4 operations case reports around 30% improvement in repair time and more than 30% reduction in backend effort; those results came from a portfolio of scenarios run through a repeatable lifecycle, not from a single clever automation.

The transferable principle is a single sentence. An AI-agent NOC/SOC is not a technology the operator installs; it is a redistribution of decision rights between the people who own the operation, the partner who builds the capability and the function that guarantees the platform — and it becomes real the day a network engineer, an integrator developer and a platform architect share one backlog, one baseline and one definition of what “done safely” means.


At TelcoCrux, we help operators design the operating model behind an autonomous NOC/SOC — the decision rights, the risk-tiered governance, the scenario portfolio and the measurement system that turn AI agents into a controlled operational capability rather than a portfolio of pilots. If you are defining how your operations, systems and integration partners should work together to get there, let’s talk.

Telco Crux Consulting
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.