Nestack Agent Care
Industries / Healthcare / Appeals & grievances agent

Healthcare AI agent · Appeals & grievances

Appeals & Grievances AI Agent (Health Plan)

Classify each case against CMS criteria, start the right clock, assemble the evidence a reviewer needs, and draft the resolution and its notice — a qualified reviewer decides, a clinician decides anything clinical.

4–6 weeksTypical delivery
Your stackDeployment
Pre-decisionReviewer-led
Agent CareAfter launch

What this agent does

Classifies the case, not the outcome

In
01

A case arrives — a grievance, a determination or an appeal — from the channel it was filed in.

02

Normalise the filer, the date received, the issue raised and the intake channel across sources.

Reason
03

Propose the classification against the intake criteria, and screen for language that makes it expedited.

04

Apply the plan's configured deadlines, notice rules and the state clinician-review trigger as written.

05

Bind the clock start, the classification and every deadline to the case record, and mark what intake leaves ambiguous.

Decide
06

Flag a grievance that reads like it contains an appeal, and a standard request that reads like it should be expedited.

07

Route every classification, every clock and every drafted resolution to the assigned reviewer for a decision.

Out
08

Retain the intake record, the classification rationale, the evidence and the reviewer's decision.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

The agent proposes the classification and drafts the resolution; a qualified reviewer decides, a clinician decides anything clinical, and a missed deadline still forwards to the independent review entity automatically.

Example workflow

One case, intake to resolution

AgentHuman
1Case receivedWritten appeal, phone grievance, fax, portal submission or provider request
2Clock and classification assembledFiler, issue raised, classification criteria, deadline rules and expedited-screening criteria
3Resolution draftedDraft resolution, notice language, cited criteria and confidence
4Controls appliedClassification checks, clock-start checks, notice-content checks and confidence threshold
No human action required

Stages 1 to 4 run unaided, and no case is closed at any of them — the agent is classifying and drafting, and the reviewer's lane opens at the confidence gate.

5DecisionBranches at the confidence threshold
High confidence

Goes to the assigned reviewer to decide.

Low confidence

Adds a compliance-officer read first.

Reviewer decision

The case is held with its classification, the evidence assembled and the confidence.

Confirm · Reclassify · Escalate
Confirmed — resolution released
6Case-management system updatedOnly where write access and approval policy allow it
7Outcome evaluatedClassification accuracy, clock adherence, reviewer edits and IRE-referral outcomes
Reviewer edits

Every reviewer edit made at decision is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Issuing, denying or closing a determination.
Deciding anything clinical.
Making the adverse-determination call reserved for a licensed clinician.
Overriding a classification once a reviewer has confirmed it.
Automation boundaryAgent acts unaided
Classify the case against CMS-defined criteria.
Start and track the clock from the date the case was received.
Assemble the evidence record — the file, the criteria.
Draft the resolution and its notice, and hold both for the reviewer's decision.
Any write happens inside the boundaries agreed at implementation, never ahead of the reviewer's decision.
Treating a missed deadline as a delay rather than an outcome.
Assuming one national rule where the clinician-review trigger varies by state.
Merging a grievance and an appeal into a single workflow.
Changes to classification criteria or clock configuration.

Example output

One case, annotated

Everything the agent drafts is attached to the case it was drawn from.

Case output · single caseIllustrative example
Case
Drafted resolution
Case type
Filing channel
Confidence
Reviewer status
Part B drug reconsideration
Draft resolution: partial reversal — reverses the Part B drug denial on the submitted clinical evidence
24-hr expedited
Filed by the treating provider
89%
Licensed clinician review confirmed
As receivedTaken from the case file and the evidence submitted — nothing on this side is decided by the agent.
Evidence assembled Provider clinical letter Denial notice and criteria Receipt date, clock start
Why this reads as reversibleThe submitted evidence meets the criterion the original denial cited.
ActionConfirmReclassifyEscalate
What the score decidesBelow the configured threshold the case picks up a compliance-officer read before.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every case filedFrom the intake channel
03Classification

Classify against criteria

Weigh the language filed against CMS's criteria for grievance, organization determination and appeal.

01Approved path

The clock is the determination

The word a filer used doesn't set the clock — the classification against CMS's criteria does.

02Human review

Point the reviewer at what needs

Cases near the expedited boundary and low-confidence classifications are marked, so the reviewer's read starts where the clock is tightest.

04Build an evidence trail

The case, the clock it was filed against and the reviewer who decided stay on the file.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Case-management systemsMHK · Casetrakker
Facets · QNXT
Clinical-criteria platformsInterQual · MCG
UM and clinical-review systems
Correspondence & noticeNotice-template engines
Member correspondence platforms

Agent

Appeals & grievances resolution

Reads the case
Classifies and drafts
Holds for decision

Audit & reporting systemsCase-tracking and workflow tools
ODAG/CDAG audit-universe systems
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the notice

The controls sit one inside the next. What they let through is named in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeNarrow the agent back to intake and routing when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, classification-rule and clock-configuration changes.Track
L4TraceabilityRecord the intake, the classification, the clock and the reviewer's decision.Record
L3Reviewer decisionHold the case for the assigned reviewer; it governs release, not whether the classification is right.Gate
L2Policy guardrailsTest the classification and the draft notice against CMS criteria and content rules; a failure returns the case.Restrict
L1Confidence thresholdsRoute low-confidence classifications to a compliance-officer read before the reviewer sees them.Require review
Model coreCase classified — draft resolution, cited criteria and confidence
L1 – L2Test whether a classification may stand
L3Puts the release in a reviewer's hands
L4 – L5Keep the case and the clock behind it
L6Reverts to intake and routing when signals degrade

How Nestack evaluates it

Evaluate the case workflow — not only the final resolution.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the resolution the member sees
Depth of coverage ▼
E1Final-output evaluationDid the classification match CMS's criteria for the case actually filed?
E2Step-level evaluationDid the agent use the right case type, clock rules and state-clinician configuration?
E3Tool evaluationDid it read and write the correct case, clock and case-file record?
E4Confidence calibrationDo low-confidence classifications actually draw more reviewer edits?
E5Slice evaluationHow does performance change across specific case types?
E6Business outcomeHow many cases needed a reviewer edit or had an incomplete record found later?
Floor — the outcome the plan answers for

Failure modes

Where each failure originates in the agent

Seven failure modes, placed at the stage each one begins.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
GV-03

Late evidence not retrieved

Supplemental clinical evidence received near the deadline never enters the file the reviewer sees.

Stage gathersCase record, prior submissions and notices
02 · Reasoning2 modes
GV-04

Keyword-driven classification

A complaint is classified as an appeal because the transcript contains the word denied.

GV-06

Clock started at assignment

The deadline is calculated from the date a case was assigned, not the date it was received.

Stage proposesThe draft resolution, the clock and confidence
03 · Tool / write2 modes
GV-02

Wrong clock applied at write

A Part B drug case writes on the 72-hour clock instead of the non-extendable 24-hour clock.

GV-05

Grievance routed as an appeal

A grievance writes into the appeal workflow, forwarding it to the IRE though it was never eligible.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
GV-01

Notice missing the cited criterion

The draft states 'not medically necessary' without naming the guideline the denial relies on.

Stage returnsThe resolution the reviewer signs
05 · Change / Version1 mode
GV-07

Silent state-rule drift

A new state's clinician-review statute takes effect before the agent's configuration requires it.

Stage tracksModel, prompt, classification and state rules
Sev-1 · issues outside the boundary Sev-2 · misclassified case, wrong clock Sev-3 · evidence degrades, goes to review

Affected slices

A blended turnaround hides the case that missed it

A blended turnaround time can look compliant while a handful of case types absorb nearly all of the misses. Nestack reports the rate by case type, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Cases arriving mid-clock from another channel7.5%3.9× Review
Expedited requests filed as standard5.6%2.9× Review
Grievances containing an appeal3.6%1.9× Watch
Standard post-service appeals1.9%0.8× Normal
Bar: misclassification rate lift vs. the standard post-service baseline · scale 0–4.0× · tick at 2.0× 2 of 4 slices over threshold

Evidence-linked improvement

The loop closes on the case, not the count

A cycle ends when the misclassified case is one the next release has to catch. That suite is what the next case opened is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Misclassification rate rises in a case-type slice.

02Diagnose

The case that aged into a determination is traced to the classification, the clock start or the queue behind it.

03Improve

Version-stamp the correction and attach the cases that exposed it.

04Verify

The affected cases run again, and a fail stops the release.

05Learn

It becomes a permanent test, and the classification rules move with it.

Learn → DetectThe return edge. The next case classified runs against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, case workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Case-workflow discovery and boundary definition.
02Case-system and channel assessment.
03Classification and clock-rule mapping and rule mapping.
04Case ingestion and normalisation.
05Classification and clock-start logic.
06Confidence scoring and flag routing.
07Reviewer decision workflow.
08Case and correspondence integration.
09Classification and clock cases.
10Guardrails and review controls.
11Case-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne case type, one plan ProductionProduction case-system integration AdvancedMultiple lines / states
Introduced at Pilot
Classification and clock recommendations
Reviewer decision
Classification-accuracy baseline
Introduced at Production
Reporting by case type
Decision workflow in your systems
Approved case write-back
Case-system integration
Introduced at Advanced
Multi-state clinician rules
Multi-stage reviewer chains
High case volume
Multi-line appeals controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your case-management system access and intake channels Case ingestion and channel mappingWeek 1
02Representative cases across case-type cohorts Classification baseline and clock-start logicWeek 2
03Your classification criteria and state clinician-review rules Classification-criteria and clock-rule mappingWeek 1
04Access to relevant APIs, feeds or exports Case-system and channel assessment, then integration setupWeek 2
05Decisions you would not want issued Deadline cases and the evaluation suiteWeek 4
06What no case may be closed without Confidence scoring, flag routing, guardrails and review controlsWeek 3
07Named reviewers and clinicians to decide cases Reviewer decision workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

The bands follow the work as it actually falls, so week 5 holds evaluation and pilot at once.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Case workflow discovery, classification mapping and boundary W2Case-system integration and the classification baseline W3Classification logic, clock rules and review controls W4Evaluation suite, guardrails and failure-mode testing W5Correspondence integration, pilot cases and corrections W6One case cycle worked under the appeals manager, then Agent Care handover
Reading the bandA bar covers the weeks its work is named in, and nothing else. The week 5 overlap is real, not padding.
At the end of W6Once the cycle validates, monitoring moves to Agent Care.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Healthcare AI agent

Build an appeals & grievances agent around your case clock.

Show us your case types, your clocks and who signs the notice. We'll map the classification workflow around your reviewer — and the clinician beside them who makes any clinical call.

Nestack Agents · Appeals & grievancesAGT-HC-17 · Agent Care available after launch