Nestack Agent Care
Industries / Healthcare / UM review agent

Healthcare AI agent · Utilization management

Utilization-Management & Medical-Necessity Review AI Agent

Locate the criteria that apply, find where the submitted record speaks to each one, and hand a nurse reviewer or medical director a structured summary — the reviewer decides, and the agent never denies.

4–6 weeksTypical delivery
Your stackDeployment
Clinician onlyDetermination
Agent CareAfter launch

What this agent does

Prepares the review, never the determination

In
01

Take the request as submitted — member, plan, service, level of care, dates and the requesting provider's own rationale.

02

Pin the criteria set, edition and plan terms that govern that request on that date of service, and stamp what was read.

Reason
03

Find the notes, results, functional scores and prior therapy that speak to each criterion, and cite each one.

04

Mark every criterion evidenced, not evidenced or not applicable, and carry the circumstances it has no field for.

05

Check that the level of care, the dates and the service being compared are the ones the request actually asks for.

Decide
06

Where every criterion is evidenced, the case goes to a reviewer as a complete file, with nothing left to chase.

07

Everything else goes to a qualified clinical reviewer, and to a medical director of the specialty the plan's policy requires.

Out
08

Hand over a structured summary — criteria set and edition, each criterion with its record location, gaps named.

09

Retain the criteria version, the record read, the evidence cited, the decision and the reviewer's own words.

Product statement

The agent finds the evidence and sets it beside the criteria. The determination is the reviewer's, and an adverse determination is issued only by a qualified clinician — never by this agent, and never on its recommendation.

Example workflow

One review, end to end

AgentHuman
1Request receivedA prior-authorisation, concurrent-review or continued-stay request with the record the provider submitted
2Criteria and plan terms pinnedThe criteria set, edition and plan document in force for that member, service and date of service, with the date read
3Evidence located in the recordEach criterion set beside the note, result, functional score or documented prior therapy that speaks to it
4Summary assembledCriterion by criterion, evidenced or not, with the page it came from and the circumstances no criterion asks about
No human action required

Stages 1 to 4 run without a person in the loop — the criteria lookup, the record search and the summary are finished before a reviewer opens the case. Nothing has been decided.

5DecisionSplits on whether every criterion is evidenced
Every criterion evidenced

Goes to a reviewer as a complete file, undecided.

Anything unevidenced or unclear

Goes to a clinical reviewer, undecided.

Nurse reviewer or medical director

Reads the record itself, decides, and writes the rationale in their own words. An adverse determination is theirs alone.

Approve · Escalate · Send back
Decided — handed back
6Determination recordedThe reviewer's decision and their own words, recorded under their name with the criteria version and evidence
7Outcome evaluatedRetrieval recall, criteria-version correctness, reviewer agreement, read-through and overturn on appeal, by payer class and service line
Reviewer changes

Every criterion a reviewer reads differently is counted, in both directions.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Deciding that a service is not medically necessary.
Issuing any adverse determination, in whole or in part.
Recommending a denial, in words or by score.
Ending, shortening or downgrading an authorised stay.
Automation boundaryAgent acts unaided
Pin the criteria set, edition and plan terms to the date of service.
Retrieve only the record the criteria call for, and cite where each finding sits.
Mark each criterion evidenced or not evidenced, and draw no conclusion from it.
Assemble the peer-to-peer and appeal packet from the file as it already stands.
Write actions run only inside the approval boundaries agreed during implementation. An adverse determination is not one of them, on any tier, in any configuration.
Choosing which criteria set or edition applies.
Writing the reviewer's clinical rationale for them.
Closing a peer-to-peer, an appeal or a grievance.
Changing criteria sources, thresholds or approval rules.

Example output

One review request, annotated

Everything the agent marks is attached to the criterion and the record line it came from.

Utilisation-review output · single requestIllustrative example
Request
Plan
Criteria edition
Criteria evidenced
Confidence
Determination
Inpatient continued stay
Managed Medicaid
In force, dated
Six of eight
91%
None — reviewer's
As receivedThe request, the plan terms and the criteria edition in force on that date of service, with the date it was read.
Evidence used Progress note, day 3 Functional score on file Documented prior therapy
Why two are openTwo criteria have no supporting text in the record. That is a statement about the record, not the patient.
ActionApproveEscalateSend back
What the score decidesHow hard the reviewer should look for evidence the agent may have missed. It does not decide the case.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every review requestFrom the UM queue and provider intake
03Criteria & plan terms

Apply the criteria that actually govern

The criteria set, edition and plan terms in force for that member, service and date of service — read at review time and stamped with the date.

01Approved path

Take the search off the reviewer

The record is read, the criteria are matched and each finding is cited before a reviewer opens the case, so the reviewer reads rather than hunts.

02Human review

Show the gaps as gaps

A criterion with nothing behind it in the record is named unevidenced and put in front of a clinician — never scored, ranked or handed over as a reason to deny.

04Build an evidence trail

Retain the criteria version, the record read, the evidence cited, the reviewer's own rationale and the appeal outcome — on both paths.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

UM platform & case fileEpic Payer Platform · GuidingCare
Jiva · UM case and queue APIs
Criteria & medical policyInterQual · MCG
Medicare NCD and LCD · plan medical policy
Clinical record exchangeEpic · Oracle Health · C-CDA and FHIR
Provider portal · attachments

Agent

Utilisation & medical-necessity review

Pins criteria
Locates evidence
Hands to reviewer

Claims, benefits & appealsCore claims · eligibility and benefits
Appeals and grievances · member notices
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and a determination

Each control wraps the one inside it. A summary clears every layer before a reviewer opens it, and the determination itself sits outside all six.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeReturn summarising to staff if evaluations or appeal signals degrade.Roll back
L5TraceabilityRecord criteria version, evidence cited, reviewer rationale and outcome.Record
L4Clinician gateA licensed clinician decides, specialty-matched where the plan requires it.Gate
L3No-determination ruleNo determination, no denial recommendation, no score that stands for one.Withhold
L2Evidence citationNo criterion is marked without a record location behind it.Cite
L1Criteria versioningCriteria set, edition and plan terms pinned to the date of service.Pin
Model coreSummary — criteria set and edition, each criterion with the record evidence behind it, and retrieval confidence
L1 – L2Decide whether the summary may stand
L3Keeps a determination out of the agent
L4Puts the decision with a licensed clinician
L5 – L6Keep the trail and pull automation back

How Nestack evaluates it

Evaluate the summary — and what the reviewer did with it.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the summary a reviewer opens
Depth of coverage ▼
E1Final-output evaluationDid it find every piece of evidence the submitted record actually holds?
E2Step-level evaluationWas the criteria set, edition and plan term right for that date of service?
E3Tool evaluationDid it read the correct member, request, encounter and record?
E4Reviewer agreementDo nurse reviewers and medical directors agree with what it marked?
E5Slice evaluationHow does performance change across payer classes, service lines and settings?
E6Business outcomeOverturn rate on appeal, and whether reviewers still open the record.
Floor — the determination that survives appeal

Failure modes

Where each failure originates in the agent

Seven failure modes plotted against the five stages of the agent lifecycle.

Agent lifecycleDirection of processing →
01 · Criteria lookup2 modes
UM-01

Wrong criteria edition

Matched against a set the plan no longer applies.

UM-02

Wrong level of care compared

Inpatient criteria run against a request for observation.

Stage pinsCriteria set, edition, plan terms and date of service
02 · Record retrieval1 mode
UM-03

Evidence present, not found

The note supports the criterion, the summary does not.

Stage gathersThe submitted notes, results, scores and prior therapy
03 · Criteria matching1 mode
UM-04

Circumstances flattened

What the checklist has no field for never reaches review.

Stage marksEach criterion evidenced, unevidenced or not applicable
04 · Summary / handover2 modes
UM-05

Summary reads as a verdict

Gaps ranked or worded so the reviewer hears a denial.

UM-06

Rationale not the reviewer's

The recorded reasoning is the summary's, not theirs.

Stage hands overThe summary a nurse reviewer or director opens
05 · Drift / Version1 mode
UM-07

Read-through decays

The summary is good enough that the record goes unopened.

Stage tracksModel, criteria editions and how reviewers use it
Sev-1 · the agent has shaped a determination Sev-2 · the file overstates the review Sev-3 · the wrong standard runs, review catches it

Affected slices

The misses concentrate where members appeal least

Aggregate retrieval recall and reviewer agreement can look acceptable while a few payer-class and service-line cohorts carry most of the missed evidence — and those are the cohorts least likely to appeal anything. Nestack reports performance by slice, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Medicaid and dual-eligible members5.9%3.6× Review
Behavioural-health requests4.6%2.8× Review
Interpreter-required records3.2%1.9× Watch
Repeat in-network requests1.2%0.7× Normal
Bar: missed-evidence lift vs. routine in-network baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

An overturn on appeal is a finding about the summary

When an appeal goes the member's way, the evidence was usually in the record that day. That belongs in the evaluation of the summary handed over.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Retrieval recall, reviewer agreement or overturn rate moves in one slice.

02Diagnose

Traced to the criteria edition, the retrieval, the marking or the wording.

03Improve

The source, retrieval or summary format is changed, re-approved by the medical director and version-linked.

04Verify

Re-run against held-out cases, including every one overturned on appeal.

05Learn

The overturned case becomes a regression case with the missed evidence marked.

Learn → DetectThe return edge. Every cycle also checks what reviewers actually opened — a summary good enough to stop them reading the record is a defect, not a result.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, criteria and plan-term sourcing, retrieval and summarisation, evaluation, reviewer workflow, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Workflow discovery and automation-boundary definition.
02Criteria, edition and plan-term sourcing.
03UM platform and record-exchange APIs.
04Date-of-service criteria and version pinning.
05Minimum-necessary retrieval scoping.
06Evidence location and citation per criterion.
07Summary format sign-off with clinicians.
08Approval-path rules and the no-determination constraint.
09Nurse and medical-director review workflow.
10Recall, agreement and equity-slice evaluation.
11Peer-to-peer and appeal packet assembly.
12Observability, deployment and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne plan, one service line ProductionProduction UM integration AdvancedMulti-plan / multi-line
Introduced at Pilot
Criteria and plan-term lookup
Evidence location and citation
Structured reviewer summary
Clinician determination gate
Baseline evaluation
Introduced at Production
Date-of-service version pinning
Concurrent and continued-stay review
Reviewer agreement and read-through tracking
Observability and audit trail
Introduced at Advanced
Peer-to-peer and appeal preparation
Multi-plan and enterprise controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on the plans and service lines in scope, criteria licences and editions, UM platform and record-exchange integrations, review volume, reviewer workflow controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01The plans, service lines and review types in scope Workflow discovery and automation-boundary definitionWeek 1
02Your criteria licences and the editions you apply Criteria, edition and plan-term sourcing, pinned to date of serviceWeek 2
03Access to the UM platform and clinical-record exchange Epic Payer Platform or GuidingCare, and C-CDA record exchangeWeek 2
04Minimum-necessary and release-of-information rules Record retrieval scoping and disclosure limitsWeek 3
05Real cases, including the ones overturned on appeal Evaluation suite, regression cases and failure-mode testingWeek 4
06Named nurse reviewers and a medical director Reviewer workflow, summary format and agreement trackingWeek 4
07Your written policy on who may issue a denial Approval-path rules, the clinician gate and the audit trailWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Phases are drawn over the weeks they actually occupy. The overlap in week 5 is where reviewer agreement is measured on the first live summaries.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Plans and review types in scope; who may issue a denial W2Criteria licences, editions and plan terms pinned to date of service W3Record retrieval, evidence citation and the summary format W4Reviewer workflow, evaluation suite, equity slices and failure-mode testing W5First live summaries, reviewer-agreement measurement and corrections W6Reviewers decide live cases from agent summaries, then handover
Reading the bandReviewer agreement in week 5 is measured on live cases, not a retrospective sample. Each bar covers its own weeks only.
At the end of W6Reviewers have decided live cases from agent summaries, and every adverse determination in that period was issued by a clinician.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Healthcare AI agent

Build a utilisation-review agent around the reviewers you already have.

Show us the criteria sets you licence, how a request reaches a reviewer, and who may issue an adverse determination. We'll assemble one case as a reviewer would want it, and agree what it may never say.

Nestack Agents · Utilisation management & medical-necessity reviewAGT-HC-11 · Agent Care available after launch