Nestack Agent Care
Industries / Telecom / NOC incident copilot

Telecom AI agent · NOC incident

NOC Incident & Root-Cause Copilot AI Agent

Correlate alarms and telemetry into one incident, rank the hypotheses that could explain it with the evidence behind each, and leave the declaration of cause to the named duty engineer.

4–6 weeksTypical delivery
Your stackDeployment
Hypothesis-onlyEngineer sign-off
Agent CareAfter launch

What this agent does

Assembles the case, not the conclusion

In
01

Ingesting alarms, counters, topology and change records from supported assurance or ticketing sources.

02

Normalising element names, timestamps and severities, and carrying each signal forward with its source.

Reason
03

Grouping related alarms into one incident, and stating what the grouping is based on.

04

Applying the correlation, topology and maintenance-window rules configured for the network.

05

Assembling the evidence pack behind each hypothesis — topology, dependency and timing.

Decide
06

Ranking hypotheses with what would falsify each one, rather than naming a single cause.

07

Surfacing the reportability question, planned maintenance included, to the duty engineer.

Out
08

Retaining the correlation, the signals it set aside, the evidence and the engineer's decision.

09

Executing write actions only inside the approval boundaries agreed during implementation.

Product statement

The notification clock runs from discovery, not from diagnosis; the copilot proposes and a duty engineer decides.

Example workflow

One incident, alarm to declaration

AgentHuman
1Alarm burst receivedFault management, performance telemetry, change record or trouble ticket
2Evidence assembledTopology, dependency, timing and recent changes, each with its named source
3Hypotheses rankedCauses, evidence, falsifiers and confidence
4Controls appliedGrouping limits, 911 and 988 path exclusions, maintenance-window checks and confidence threshold
No human action required

Stages 1 to 4 run unaided, and nothing is declared or closed at any of them — the copilot is assembling, and the engineer's lane opens at the confidence gate.

5DecisionBranches at the confidence threshold
High confidence

Goes to the duty engineer to accept.

Low confidence

Adds a senior NOC read first.

Engineer review

The incident is held with its evidence, its ranked hypotheses and the confidence.

Accept · Revise · Send to senior review
Accepted — cause declared
6Incident record updatedOnly where write access and approval policy allow it
7Outcome evaluatedRank at acceptance, revisions, set-aside-signal outcomes and post-incident corrections
Revisions

Every engineer revision is counted.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Declaring the root cause of an incident as a finding.
Closing an incident or declaring service restored.
Deciding whether an outage or event is reportable.
Notifying a PSAP, the FCC or any other regulator.
Automation boundaryAgent acts unaided
Correlate alarms and telemetry into a single incident view.
Assemble the evidence pack behind each candidate hypothesis.
Rank the hypotheses, and state what would falsify each one.
Surface the reportability question and its clock to an engineer.
Writes stay inside the agreed boundaries and never reach a network element — that belongs to remediation.
Changing any network element or its configuration.
Suppressing any signal on a 911 or 988 call path.
Writing the best-known cause into a notification.
Changes to correlation, suppression or escalation rules.

Example output

One incident, annotated

Everything the copilot proposes is attached to the signals it was drawn from.

Root-cause output · single incidentIllustrative example
Incident
Leading hypothesis
Alarm group
Evidence class
Confidence
Escalation
Regional transport
Protection switch on a shared fibre span, not the access nodes that alarmed
One transport ring
Topology and timing
71%
Duty engineer, clock shown
As receivedTaken from fault management and the change record — nothing on this side is inferred by the copilot.
Evidence used Topology dependency Change record Timing sequence
Why this hypothesisThe loudest element is the most instrumented one, not the cause.
ActionAcceptReviseSend to senior review
What the score decidesBelow the configured threshold the incident picks up a senior read before.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every incidentFrom the alarm stream
03Correlation

Correlate and evidence

Draw on topology, dependency, timing and the change record for the domains in scope.

01Approved path

A cause is a hypothesis

Related alarms arrive already grouped, with the basis of the grouping stated.

02Human review

Send the read to the ranked few

Ranked hypotheses and their falsifiers are marked, and nothing that would delay a statutory notification is held back for tidiness.

04Build an evidence trail

The correlation, the evidence it rests on and the engineer who accepted it stay on the incident.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Fault and alarm managementIBM Netcool · Nokia NSP
Ericsson OSS · Huawei U2000
Performance and telemetryPrometheus · Kafka streams
Splunk · Elastic · gNMI
Topology and inventoryNetCracker · Amdocs
Blue Planet · Cisco Crosswork

Agent

NOC incident & root-cause

Reads the signals
Ranks the hypotheses
Holds for the engineer

Incident and changeServiceNow · Jira
PagerDuty · change records
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the incident

The layers sit one inside the next. What none of them catches is in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modePull the copilot back to alarm listing when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, correlation-rule and topology-configuration changes.Track
L4TraceabilityRecord the signals, what was set aside, the ranking, the evidence and the decision.Record
L3Engineer sign-offHold the incident for the duty engineer; it governs declaration, not whether the hypothesis is right.Gate
L2Policy guardrailsTest grouping and correlation against the 911, 988 and special-facility exclusions; a failure returns the incident.Restrict
L1Confidence thresholdsRoute low-confidence rankings to a senior NOC read before the duty engineer sees them.Require review
Model coreHypotheses ranked — candidate causes, evidence, falsifiers and confidence
L1 – L2Test whether a hypothesis may stand
L3Puts the declaration in an engineer's hands
L4 – L5Hold the evidence the hypothesis rests on
L6Falls back to alarm listing when signals degrade

How Nestack evaluates it

Evaluate the whole triage path — not only the cause that was accepted.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the cause the incident record carries
Depth of coverage ▼
E1Final-output evaluationDid the accepted cause survive the post-incident review?
E2Step-level evaluationDid the copilot use the right topology, change record and correlation rules?
E3Tool evaluationDid it read the correct domain and write the correct incident?
E4Confidence calibrationDo low-confidence rankings actually attract more engineer revisions?
E5Slice evaluationHow does performance change across specific incident classes?
E6Business outcomeHow many incidents needed a revision or a correction after closure?
Floor — the outcome the carrier answers for

Failure modes

Where each failure originates in the agent

Seven failure modes, placed at the stage each one originates.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
GK-03

Stale topology view

Dependency read from a superseded topology record.

Stage gathersAlarms, telemetry, topology , with the source each came from
02 · Reasoning2 modes
GK-04

Correlation read as causation

The most instrumented element is ranked as the cause.

GK-06

Cascade masks the trigger

A loud downstream storm outranks the quiet change behind it.

Stage proposesCandidate causes, evidence, falsifiers and confidence
03 · Tool / write2 modes
GK-02

Grouping on a 911 path

A signal touching an emergency path is grouped away.

GK-05

Reportability left unraised

The clock runs while no one is asked the question.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
GK-01

Cause stated as finding

A ranked hypothesis is recorded as a settled cause.

Stage returnsThe cause the engineer accepts and the record carries
05 · Change / Version1 mode
GK-07

Silent correlation regression

A model or rule change widens what the copilot will assert.

Stage tracksModel, prompt, correlation rules and topology config
Sev-1 · acts outside the boundary Sev-2 · a wrong cause reaches the record Sev-3 · signal degrades, incident routes to review

Affected slices

Overall quality can hide one bad class

Do not read the revision rate whole — break it out by incident class. No audited production figure exists for automated root-cause analysis in live networks, so Nestack reports what your engineers revise, by slice.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Alarm-storm incidents9.9%3.9× Review
Shared-infrastructure faults6.9%2.7× Review
Post-change incidents4.6%1.8× Watch
Single-element faults1.5%0.6× Normal
Bar: engineer-revision-rate lift vs. single-element baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

A post-mortem is not the end of the loop

The cycle ends in a regression case, not in a meeting about what happened. That suite is what the next incident into the NOC is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Engineer-revision rate rises in an incident class.

02Diagnose

The evidence pack behind each disputed hypothesis is pulled and read until the miss narrows to one.

03Improve

Every change is versioned against the incidents that exposed it.

04Verify

A failing case holds the release back.

05Learn

The case joins the suite for good, and the correlation rules are revisited.

Learn → DetectThe return edge. The next detection runs against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, triage workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01NOC workflow discovery and automation-boundary definition.
02Assurance and telemetry source assessment.
03Topology, dependency and correlation-rule mapping.
04Alarm ingestion and normalisation.
05Correlation and hypothesis ranking.
06Confidence scoring and escalation routing.
07Duty-engineer review workflow.
08Assurance and ticketing integration.
09Causal-claim regression cases.
10Guardrails and escalation controls.
11Incident-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne domain, one NOC ProductionProduction assurance stack AdvancedMultiple domains / regions
Introduced at Pilot
Correlation to your topology
Engineer sign-off
Hypothesis-quality baseline
Introduced at Production
Reporting by incident class
Review workflow in your systems
Approved incident write-back
Assurance-stack integration
Introduced at Advanced
Multi-vendor correlation rules
Multi-stage NOC escalation
High alarm volume
Multi-domain incident controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your alarm sources and topology model Alarm ingestion and topology mappingWeek 1
02Representative past incidents Correlation baseline and hypothesis rankingWeek 2
03Your correlation and suppression rules Correlation-rule mapping and automation-boundary definitionWeek 1
04Access to relevant APIs, feeds or exports Assurance and telemetry assessment, then integration setupWeek 2
05Causes you would not want acted on Causal cases and the evaluation suiteWeek 4
06What must reach an engineer before anything is touched Confidence scoring, escalation routing, guardrails and review controlsWeek 3
07Named duty engineers to review incidents Duty-engineer review workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

The bands follow the real work, which is why evaluation and pilot share the fifth week.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1NOC workflow discovery, rule mapping and the automation boundary W2Assurance integration and the correlation baseline W3Triage workflow, confidence logic and escalation controls W4Evaluation suite, exclusion checks and failure-mode testing W5Ticketing integration, pilot incidents and targeted corrections W6One incident cycle triaged under the NOC, then handover
Reading the bandA bar covers the weeks its work is named in, and nothing else. The week 5 overlap is real, not padding.
At the end of W6The cycle closes validation and Agent Care assumes monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Telecom AI agent

Build a NOC copilot around the engineers who declare cause.

Show us your alarm sources, your topology and who declares cause. Being wrong costs twice: the fault stays live and the clock keeps running.

Nestack Agents · NOC incident & root-causeAGT-TL-08 · Agent Care available after launch