Nestack Agent Care
Industries / Insurance & Financial Services / Legal-bill review agent

Insurance AI agent · Legal-bill review

Legal-Bill Review & Litigation-Spend AI Agent

Read each outside-counsel invoice against the litigation guidelines for that matter, flag entries that don't match, track spend against reserve, and hold every adjustment for the litigation manager who decides.

4–6 weeksTypical delivery
Your stackDeployment
Pre-adjustmentManager review
Agent CareAfter launch

What this agent does

Prepares the adjustment, not the decision to make it

In
01

An invoice arrives from outside counsel, and the agent pulls the guideline set for that matter.

02

A matter opens, and the agent maps entries, task codes and expense lines against the guideline.

Reason
03

Each entry is compared to the guideline provision it falls under, not judged on its own.

04

The guideline applies only once the matter's jurisdiction is confirmed to enforce it.

05

Every flagged entry is bound to the guideline clause and the invoice line it was read against.

Decide
06

Spend against the matter's litigation budget and reserve is tracked using the erosion rule the policy sets.

07

The flagged review is routed to the litigation manager, with any reservation-of-rights conflict escalated separately.

Out
08

The entry, the guideline it was read against, the adjustment reason and the consent record stay on the matter.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

The agent prepares the flagged review and its reason; the litigation manager decides on any adjustment, and outside counsel is told and can dispute it.

Example workflow

One invoice, receipt to adjustment

AgentHuman
1Invoice receivedOutside-counsel invoice, matter number and the guideline set that applies to it
2Context assembledTime entries, task codes, expense lines, reserve and the matter's guideline, each with a source
3Review draftedFlagged entries, proposed adjustment, reason and confidence
4Controls appliedGuideline-match checks, budget-and-reserve checks, privileged-narrative handling and confidence threshold
No human action required

Stages 1 to 4 run unaided, and no adjustment is made at any of them — the agent is drafting the review, and the manager's lane opens at the confidence gate.

5DecisionBranches at the confidence threshold
High confidence

Goes to the litigation manager to decide.

Low confidence

Adds a legal-ops read before the decision.

Manager review

The review is held with its flagged entries, the guideline and the confidence.

Approve · Edit · Escalate for conflict review
Approved — counsel notified
6Matter record updatedOnly where write access and approval policy allow it
7Outcome evaluatedAdjustment accuracy, dispute outcomes, reserve movement and post-adjustment corrections
Edits

Every manager edit before the decision is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Adjusting, reducing or rejecting an invoice line.
Withholding or releasing payment to outside counsel.
Setting or approving the matter's litigation budget.
Instructing counsel on strategy, staffing or how to litigate the matter.
Automation boundaryAgent acts unaided
Read each invoice against the guideline that applies.
Track spend against the matter's litigation budget and reserve.
Flag entries that do not match the applicable guideline provision.
Prepare the adjustment and its reason, and hold it.
Any write happens inside the boundaries agreed at implementation, never ahead of the decision.
Deciding a coverage or reservation-of-rights question.
Deciding whether independent counsel is owed.
Sending detailed invoices to a bill-review vendor absent documented insured consent.
Changes to the guideline library or adjustment thresholds.

Example output

One entry in the invoice, annotated

Everything the agent flags is attached to the guideline it was read against.

Invoice-review output · single invoiceIllustrative example
Matter
Flagged entry
Stated amount
Guideline clause
Confidence
Attribution
Products-liability defence, RoR matter
Two associates billed the same deposition prep on the same day
$1,240.00
Block-billing guideline
89%
Litigation manager of record
As receivedTaken from the invoice and the guideline citation given — nothing on this side is written by the agent.
Guidelines applied Block-billing clause Task-code requirement Prior invoice history
Why this flagIt states what the guideline requires, not a judgment on the work.
ActionApproveEditEscalate for conflict review
What the score decidesBelow the configured threshold, the review picks up a legal-ops read before the manager sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every invoiceFrom the matter record
03Review

Read against the guideline

Draw on the guideline that applies and the matter's configured budget and reserve rules.

01Approved path

The insured is the client

Routine entries that match the guideline are cleared without a flag.

02Human review

Point the manager at what

Entries that don't match the guideline are flagged, so review starts where spend is at risk.

04Build an evidence trail

The entry, the guideline it was read against and the manager who adjusted it stay on the matter.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

E-billing & matter systemsLegal Tracker · CounselLink
Collaborati · Passport
Guideline & budget dataLitigation guidelines
Reserve and budget records
Claims & policy systemsClaim file · reserve ledger
Guidewire · Duck Creek

Agent

Legal-bill review & litigation spend

Reads the guideline
Flags the entry
Holds for the manager

Counsel correspondenceDispute tracking
Adjustment correspondence
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the invoice

Each layer contains the next. What escapes all of them is named in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeFall back to flagging only when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, guideline and reserve-configuration changes.Track
L4TraceabilityRecord the entry, the guideline it was read against, the flag and the manager's decision.Record
L3Manager reviewHold the adjustment for the litigation manager; it governs release, not whether the flag is right.Gate
L2Policy guardrailsTest the flag against the matter's guideline and privilege-handling rules; a failure returns it.Restrict
L1Confidence thresholdsRoute low-confidence flags to a legal-ops read before the manager sees them.Require review
Model coreReview produced — flagged entries, proposed adjustment, reason and confidence
L1 – L2Test whether the flag may stand
L3Puts the adjustment in the manager's hands
L4 – L5Keep the entry and the guideline behind it
L6Falls back to budget reporting when signals degrade

How Nestack evaluates it

Evaluate the review workflow — not only the finished adjustment.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the adjustment counsel sees
Depth of coverage ▼
E1Final-output evaluationDid every flagged entry match the guideline clause it was read against?
E2Step-level evaluationDid the agent use the right guideline, budget and reserve configuration?
E3Tool evaluationDid it read and write the correct matter and the correct invoice field?
E4Confidence calibrationDo low-confidence flags actually attract more manager edits?
E5Slice evaluationHow does performance change across specific matter types?
E6Business outcomeHow many flags needed a manager edit or drew a dispute from counsel?
Floor — the outcome the carrier answers for

Failure modes

Where each failure originates in the agent

Seven failure modes, placed at the stage each one originates.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
LB-03

Superseded guideline version

The matter's litigation guideline is read from a version the carrier has since revised.

Stage gathersInvoice, matter number, guideline and reserve
02 · Reasoning2 modes
LB-04

Guideline misapplied

A guideline is applied in a state where it may not be enforceable against counsel.

LB-06

Adjustment overstated

A proposed adjustment is presented as decided, not as a flag for the manager.

Stage proposesFlagged entries, adjustment and confidence
03 · Tool / write2 modes
LB-02

Confidence dropped in transit

A flagged entry reaches the manager without its confidence attached.

LB-05

Duplicate invoice processed

The same invoice is reviewed twice under two different matter numbers.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
LB-01

Privilege flag dropped

A narrative entry carrying privileged detail is surfaced without the privilege marker.

Stage returnsThe adjustment the manager decides and counsel sees
05 · Change / Version1 mode
LB-07

Silent guideline drift

A guideline or reserve-configuration change widens what the agent will flag unreviewed.

Stage tracksModel, prompt, guideline and reserve configuration
Sev-1 · adjustment made outside the boundary Sev-2 · counsel not told of an adjustment Sev-3 · guideline degrades, routes to review

Affected slices

A clean docket can hide where spend concentrates

A portfolio adjustment rate can look controlled while a handful of firms or matter types carry most of the flagged spend. Nestack reports the adjustment-dispute rate by slice, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Matters under a reservation of rights7.8%4.0× Review
Block-billed and grouped time entries5.3%2.7× Review
Expense and vendor pass-through lines3.5%1.8× Watch
Task-coded time inside the guidelines1.9%0.8× Normal
Bar: adjustment-dispute rate lift vs. the task-coded baseline · scale 0–4.0× · tick at 2.0× 2 of 4 slices over threshold

Evidence-linked improvement

Every cycle closes on one more matter

A cycle is closed when the wrong adjustment is a case the next release must survive. That suite is what the next invoice reviewed is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Adjustment-dispute rate rises in a matter slice.

02Diagnose

The line that was cut without anyone telling counsel is traced back through the guideline and the invoice until the cause narrows to one entry.

03Improve

The change ships against a version, with the matters that exposed it attached.

04Verify

Release is blocked until the affected adjustment cases pass again.

05Learn

The case joins the standing suite and the guideline mapping moves with it.

Learn → DetectThe return edge. Detection next time runs against a suite one matter longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, guideline sources, review workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Review workflow discovery and scope and boundary definition.
02Matter and e-billing system assessment.
03Guideline library and reserve mapping and rule mapping.
04Invoice ingestion and normalisation.
05Flagging logic and guideline binding.
06Confidence scoring and escalation routing.
07Litigation-manager review workflow.
08E-billing and claims-system integration.
09Adjustment and privilege cases.
10Guardrails and dispute controls.
11Matter-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne guideline set, one matter ProductionProduction e-billing systems AdvancedMultiple guideline sets / firms
Introduced at Pilot
Flagging to your guideline and reserve rules
Manager review
Adjustment-accuracy baseline
Introduced at Production
Reporting by firm
Manager workflow in your systems
Approved write-back
Matter-system integration
Introduced at Advanced
Multi-jurisdiction guideline complexity
Multi-stage litigation approvals
High invoice volume
Multi-jurisdiction defence controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, invoice volume, review controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your litigation guideline library and reserve rules Guideline ingestion and matter mappingWeek 1
02Representative past invoices, including disputes Flagging baseline and guideline-binding baselineWeek 2
03Your budget thresholds and reserve rules Guideline library and reserve mappingWeek 1
04Access to relevant APIs, feeds or exports E-billing and claims-system assessment, then integration setupWeek 2
05Adjustments you would not want made Guideline cases and failure-mode testingWeek 4
06What no adjustment may be made without Confidence scoring, escalation routing, guardrails and dispute controlsWeek 3
07Named litigation managers to review flags Manager review workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Each phase sits over the weeks it really occupies, and week 5 carries evaluation alongside the pilot.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Review workflow discovery, guideline mapping and the automation boundary W2E-billing and claims-system integration and the flagging baseline W3Flagging workflow, confidence logic and manager-review controls W4Evaluation suite, privilege-handling checks and failure-mode testing W5E-billing integration, pilot invoices and targeted corrections W6One billing cycle reviewed under the litigation manager, then handover
Reading the bandA bar covers the weeks its work is named in, and nothing else. The week 5 overlap is real, not padding.
At the end of W6The last checks clear on live invoices and Agent Care picks up monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Insurance AI agent

Build a legal-bill review agent around your litigation guidelines.

Show us your litigation guidelines, your reserve rules and who decides an adjustment. Not a cost-cutting tool. A guideline match held for the litigation manager, with counsel told and able to dispute it.

Nestack Agents · Legal-bill review & litigation spendAGT-INS-18 · Agent Care available after launch