Nestack Agent Care
Industries / Telecom / Remediation agent

Telecom AI agent · Network remediation

Self-Healing Network Remediation AI Agent

Detect the fault, diagnose it within the configured scope, and propose a change with its blast radius and its reversion plan — then hold, because applying it is a network engineer's decision.

4–6 weeksTypical delivery
Your stackDeployment
Peer-reviewedChange approval
Agent CareAfter launch

What this agent does

Proposes the change, does not apply it

In
01

Telemetry, alarms and configuration state arrive from supported assurance and orchestration.

02

Element names, alarm codes and timestamps are normalised, and each signal keeps the path it arrived on.

Reason
03

A fault hypothesis is formed inside the domain and the elements the agent has been scoped to.

04

The carrier's change rules, maintenance windows and freeze periods are applied to whatever is proposed.

05

Blast radius is computed for the proposed change — the elements, services and paths it would touch.

Decide
06

911, 988 and special-facility paths sit outside the action space, as an exclusion rather than a setting.

07

The proposal is routed to a named network engineer for peer review and authorisation.

Out
08

The state it read, the change it wrote, the review and the authorisation are retained against the incident.

09

Write actions run only inside the approval boundaries agreed during implementation.

Product statement

The agent proposes a change; a network engineer peer-reviews and authorises it, and the carrier stays the party answerable to its regulator.

Example workflow

One fault, alarm to authorised change

AgentHuman
1Fault detectedAssurance alarms, performance counters, probes or ticket intake
2State assembledTopology, running configuration, recent changes and the freeze calendar, each with its source
3Change proposedChange, blast radius, reversion, confidence
4Controls appliedScope checks, the 911-path exclusion, the blast-radius limit, freeze-window checks and confidence threshold
No human action required

Stages 1 to 4 run unaided and change nothing — the agent is proposing, and the engineer's lane opens at the confidence gate.

5DecisionBranches at the confidence threshold
High confidence

Goes to the duty network engineer to authorise.

Low confidence

Adds a second engineer's peer review first.

Change approval

The change is held with the state it was written against, its blast radius and the confidence.

Authorise · Amend · Send to change board
Authorised — staged for application
6Change applied in stagesOnly where write access and the change process allow it
7Outcome evaluatedService-level recovery, reversions, escalations and post-change incidents
Amendments

Every engineer amendment is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Anything on a 911, 988 or special-facility path.
Applying any change to the production network.
Any change beyond the configured blast radius.
Acting in a declared incident without the commander.
Automation boundaryAgent acts unaided
Detect a fault and correlate the alarms it raised.
Diagnose within the domain and the elements it has been scoped to.
Propose a change with its blast radius and its reversion plan.
Hold the proposal for a named engineer, with the evidence attached.
Any write happens inside the boundaries agreed at implementation, never ahead of authorisation.
Reverting a change outside the bounded scope.
Declaring the fault cleared or the service restored.
Changes crossing into a partner or third-party domain.
Changes to its own thresholds, scope or radius limits.

Example output

One proposed change, annotated

Everything the agent proposes is attached to the state it was read from.

Remediation output · single proposalIllustrative example
Element
Proposed action
Blast radius
State written against
Confidence
Reversion
Edge router, metro ring
Restore the ring-facing link weighting to its peer-reviewed value
Single metro ring
Reviewed baseline
88%
Prepared, not guaranteed
As receivedTaken from the assurance feed and the configuration baseline — nothing on this side is written by the agent.
Evidence used Alarm correlation Topology and dependency Recent change record
Why this changeIt restores a value the change record shows was reviewed — the line the engineer weighs.
ActionAuthoriseAmendSend to change board
What the score decidesBelow the configured threshold the proposal picks up a second engineer's review before.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every detected faultFrom the assurance feed
03Diagnosis

Propose against the state

Draw on topology, the configuration baseline and the change record the carrier already keeps.

01Approved path

A change is still a change

Routine remediations arrive already written, with their blast radius computed.

02Human review

Send review to the risky changes

Wide-radius and low-confidence proposals are marked, so the engineer's review starts where the exposure is.

04Build an evidence trail

The proposed change, its blast radius and the engineer who approved it stay on the record.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Assurance and fault managementIBM Netcool · Nokia NSP
Ericsson EIAP · Huawei iMaster
OrchestrationONAP · Ansible Automation
Cisco NSO · Terraform
Inventory and topologyNetBox · Nokia NIM
Cramer · Granite

Agent

Self-healing network remediation

Reads the state
Proposes the change
Holds for approval

Change and incidentServiceNow · Jira
BMC Remedy · change calendars
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the network

The controls sit inside one another. What none of them catches is set out below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modePull the agent back to detect-and-propose when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, scope and blast-radius configuration changes.Track
L4Change traceabilityRecord the state read, the proposal, the review, the authorisation and what followed.Record
L3Engineer approvalHold changes for a named engineer; it governs application, not whether the change is right.Gate
L2Blast-radius limitsTest each proposal against the configured scope, the 911 exclusion and the radius limit; a failure returns it.Restrict
L1Confidence thresholdsRoute low-confidence proposals to a second engineer before authorisation is asked for.Require review
Model coreChange proposed — action, blast radius, reversion plan and confidence
L1 – L2Test whether a change may stand
L3Puts the application in an engineer's hands
L4 – L5Keep the change and the state it was written against
L6Narrows to detect-and-propose when signals degrade

How Nestack evaluates it

Evaluate the remediation workflow — not only the change that was applied.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the change the network takes
Depth of coverage ▼
E1Final-output evaluationDid the applied change produce the service-level recovery it predicted?
E2Step-level evaluationDid the agent read the right topology, baseline and change record?
E3Tool evaluationDid it read and write the correct element and the correct parameter?
E4Confidence calibrationDo low-confidence proposals actually attract more engineer amendments?
E5Slice evaluationHow does performance change across specific fault classes?
E6Business outcomeHow many changes were reverted, escalated or followed by a second incident?
Floor — the state the network is left in

Failure modes

Where each failure originates in the agent

Seven ways a change goes wrong, placed by stage.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
YR-03

Single-pipeline telemetry

State read from one monitoring path that was itself degraded.

Stage gathersAlarms, topology, configuration baseline and change record
02 · Reasoning2 modes
YR-04

Cascade mistaken for cause

The loudest downstream symptom is diagnosed as the fault.

YR-06

Conflicting closed loops

Two concurrent loops act against each other's remedy.

Stage proposesAction, blast radius, reversion plan and confidence
03 · Tool / write2 modes
YR-02

Blast radius understated

The change reaches elements the estimate did not include.

YR-05

Reversion leaves a third state

A partly applied change is reverted into neither state.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
YR-01

Recovery load after the fix

Reversion completes while the service is still failing.

Stage returnsThe change the engineer authorises and
05 · Change / Version1 mode
YR-07

Silent scope regression

A model or rule change widens what the agent will propose.

Stage tracksModel, prompt, scope rules and radius config
Sev-1 · applied outside the boundary Sev-2 · a wrong change reaches the network Sev-3 · telemetry degrades, proposal routes to review

Affected slices

A healthy total can hide one fault class

Base rates are not shared: each fault class carries its own, and one aggregate amendment rate quietly averages them together. Nestack reports the amendment rate by slice, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Core and control-plane faults9.6%3.5× Review
Cross-domain transport faults5.8%2.1× Review
RAN cell-level degradation4.1%1.5× Watch
Single-element access faults2.5%0.9× Normal
Bar: amendment-rate lift vs. single-element access baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

The loop closes on a case, not a cause

The cycle ends when a case exists in the suite, not when the miss was discussed. That suite is what the next change proposed to the network is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Amendment rate rises in a fault class.

02Diagnose

Not the loudest alarm — an engineer reads the state, the change record and the proposals until the cause narrows to one.

03Improve

Changes carry a version and the incidents that prompted them.

04Verify

Release is held until the affected cases pass.

05Learn

The case is kept permanently, and the change rules move with it.

Learn → DetectThe return edge. The next detection is measured against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, remediation workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Remediation workflow discovery and boundary definition.
02Assurance and orchestration assessment.
03Change-process, blast-radius and 911-exclusion mapping.
04Telemetry ingestion and normalisation.
05Diagnosis logic and state binding.
06Confidence scoring and proposal routing.
07Engineer approval workflow.
08Orchestration and change-system integration.
09Blast-radius and rollback cases.
10Guardrails and approval controls.
11Change-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne domain, one region ProductionProduction change process AdvancedMultiple domains / vendors
Introduced at Pilot
Proposals against your state and rules
Engineer approval
Proposal-quality baseline
Introduced at Production
Reporting by domain and element
Change-system workflow integration
Authorised change execution
Orchestration-stack integration
Introduced at Advanced
Multi-vendor change policies
Multi-stage change approvals
High change volume
Multi-domain change controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, change volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your change process and its approval chain Change-process mapping and boundary definitionWeek 1
02Representative faults and the changes that cleared them Diagnosis baseline, state binding and blast-radius modellingWeek 2
03Your topology, baseline configuration and freeze calendar Blast-radius, freeze-window and 911-exclusion mappingWeek 1
04Access to assurance, inventory and orchestration APIs Assurance and orchestration assessment, then integration setupWeek 2
05Changes you would not want applied Rollback cases and failure-mode testingWeek 4
06What a change may never touch on its own Confidence scoring, scope limits, guardrails and approval controlsWeek 3
07Named network engineers to review proposals Engineer approval workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

The bands follow the actual work, which is why the fifth week doubles rather than pads.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Remediation workflow discovery, change mapping and the automation boundary W2Assurance and orchestration integration, and the diagnosis baseline W3Proposal workflow, confidence logic and approval controls W4Evaluation suite, blast-radius checks and failure-mode testing W5Orchestration integration, pilot changes and targeted corrections W6One change window run under network engineering, then handover
Reading the bandA bar covers the weeks its work is named in, and nothing else. The week 5 overlap is real, not padding.
At the end of W6The window closes validation and Agent Care owns the running agent.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Telecom AI agent

Build a remediation agent around your change process.

Show us your assurance stack, your change process and your freeze calendar. The engineer who signs a change today still signs it after this ships — we map the boundary around that signature.

Nestack Agents · Self-healing network remediationAGT-TL-09 · Agent Care available after launch