Nestack Agent Care
Industries / Accounting / Expense-audit agent

Accounting AI agent · Expenses

Expense-Audit & Policy-Compliance AI Agent

Audit every claim against your expense policy, evidence each flag with the receipt, card line and rule behind it, and route it to a named reviewer — the agent never rejects a claim itself.

4–6 weeksTypical delivery
Your stackDeployment
Every flagHuman review
Agent CareAfter launch

What this agent does

Audits every claim, evidences every flag

In
01

Ingest expense claims, receipt images, corporate-card feeds, mileage logs and per-diem entries.

02

Normalise merchant, date, currency and category, and read the claimant's grade and entitlement.

Reason
03

Test each line against the policy version in force on the date the money was spent.

04

Read receipt, claim and card line together to find duplicates, split claims and missing evidence.

05

Check mileage, per-diem overlaps and VAT/GST reclaim eligibility against the rules for that country.

Decide
06

Raise a flag carrying the rule, the evidence and a reason written for the person being flagged.

07

Rank flags so reviewers see the defensible ones first, and pass the rest through unflagged.

Out
08

Retain the rule version, evidence, reviewer decision and the reason the reviewer gave for it.

09

Write the flag and its evidence only — clearing, rejecting and reimbursing stay with people.

Product statement

The agent flags and evidences inside the boundaries agreed during implementation; rejecting a claim, and anything that follows from it, stays with a named human.

Example workflow

One expense claim, end to end

AgentHuman
1Claim receivedExpense platform, card feed, mileage log or receipt capture
2Evidence gatheredReceipt, card line, prior claims, and the claimant's grade and entitlement
3Policy evaluatedLimits, categories, class of travel, per-diem and mileage rates
4Duplicate and pattern checksThe same spend across cycles or across card and cash, and split amounts
No human action required

Stages 1 to 4 run without a person in the loop — nothing reaches the claimant or their manager before the gate.

5DecisionSplits on the confidence threshold
High confidence

Passes with no flag raised.

Low confidence

Held as a flag for review.

Human review

The claim is held with the rule it tripped, the evidence and a reason the claimant can read.

Clear · Uphold · Ask the claimant
Decided — handed back
6Expense platform updatedThe reviewer's decision and its reason, where write access allows it
7Outcome evaluatedFlag precision, flags overturned in review, and rates by team
Overturns

Flags cleared in review are counted, by team and by grade.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Rejecting a claim or withholding reimbursement.
Calling a claim fraudulent.
Referring a person to HR or an investigation.
Granting or recording a policy exception.
Automation boundaryAgent acts unaided
Read the claim, the receipt and the matching card line together.
Test each line against the policy version in force.
Detect duplicates, split claims and per-diem overlaps.
Raise a flag carrying its rule, its evidence and a written reason.
Write actions run only inside the approval boundaries agreed during implementation. Rejecting a claim is not one of them.
Overriding a grade, per-diem or mileage entitlement.
Changing the expense policy or its thresholds.
First claims from a new joiner.
Any flag the claimant contests.

Example output

One expense claim, annotated

Every flag is attached to the line it came from and the rule it was tested against.

Audit output · single claim lineIllustrative example
Merchant
Category
Amount
Flag raised
Confidence
Routed to
Hotel, overseas trip
Accommodation
$412.00
Nightly cap exceeded
92%
The claimant's approving manager
As receivedTaken from the expense platform and the card line it settled on — nothing on this side is inferred.
Evidence used Itemised receipt Card settlement line City cap for that date
Why this flagThe city cap in the policy version in force that night — not a judgement about anyone.
ActionClearUpholdAsk the claimant
What the score decidesAbove the threshold the flag reaches a reviewer with its evidence. A human clears or upholds it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
All expense claimsFrom every submission channel
03Policy checks

Apply client-specific context

Use the policy version in force, grade entitlements and per-diem tables.

01Approved path

Check the population, not a sample

Every claim is tested against every rule, not the percentage a sampling audit draws.

02Human review

Send people only defensible flags

Reviewers get ranked, evidenced flags with a reason attached, rather than a queue of claims to read in full.

04Build an evidence trail

Retain the rule version, evidence, confidence, evaluator result and the reviewer's decision and reason — on both paths.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Expense & ERP systemsSAP Concur · Expensify · Navan
NetSuite · ERP/GL APIs
Corporate-card feedsCard transaction feeds · settlement files
Issuer APIs · statement imports
Receipt captureMobile capture · email intake
E-receipt feeds · scanned documents

Agent

Expense audit & policy compliance

Reads the evidence
Tests against policy
Evidences each flag

HR & approval workflowEmployee master · grade and entitlement
Approval systems · email
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the person flagged

Each control wraps the one inside it. A flag clears every layer before anyone is asked to look at it, and the decision sits outside all six.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeRestrict automation if evaluations or fairness signals degrade.Roll back
L5Fairness monitoringCompare flag and overturn rates across teams, grades and claim types.Compare
L4TraceabilityRecord the rule, evidence, reviewer decision and the reason given.Record
L3Human decisionDefine who reviews a flag and who may act on what it says.Gate
L2Policy guardrailsTest only against the policy version in force on the date of spend.Restrict
L1Confidence thresholdsLow-confidence checks are logged, not raised against a person.Hold back
Model coreCheck result proposed — rule tested, evidence gathered, reason drafted and confidence
L1 – L2Decide whether a flag is raised at all
L3Decides who may act on what it says
L4 – L5Keep the reason and the flag rates visible
L6Pulls automation back when signals degrade

How Nestack evaluates it

Evaluate the full workflow — not only the final flag.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the flag the reviewer and the claimant see
Depth of coverage ▼
E1Final-output evaluationWas the flag correct, and was the right rule cited?
E2Step-level evaluationDid the agent read the receipt, card line and entitlement correctly?
E3Tool evaluationDid it read or write to the correct claim and claimant?
E4Confidence calibrationDo low-confidence checks actually produce more false flags?
E5Fairness evaluationDo flag and overturn rates differ across teams and grades?
E6Business outcomeHow many flags survived review, and what reclaim was recovered?
Floor — the decision a named human signs

Failure modes

Where each failure originates in the agent

Seven failure modes plotted against the five stages of the agent lifecycle.

Agent lifecycleDirection of processing →
01 · Capture / retrieval2 modes
EX-01

Cropped receipt image

The itemisation or the tax line is cut off.

EX-02

Card and cash duplicate

The same meal claimed once on the card feed and once in cash.

Stage gathersClaim lines, receipts, card transactions and entitlements
02 · Policy evaluation2 modes
EX-03

Split claim under a limit

One dinner submitted as two lines under the cap.

EX-04

Stale policy version

Tested against today's limits, not that night's.

Stage testsRules, limits, per-diem and mileage rates, confidence
03 · Tool / write1 mode
EX-05

Flag on the wrong line

The result lands on another claim in the report.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
EX-06

Flag without its reason

The rule fires but no evidence reaches review.

Stage returnsThe flag, its evidence and the reason given
05 · Change / Version1 mode
EX-07

Flag rate drifts by team

A rule change flags one team more than the rest.

Stage tracksPolicy versions, model, prompt and rule changes
Sev-1 · a person is flagged unfairly Sev-2 · a real breach passes unflagged Sev-3 · input degrades, claim routes to review

Affected slices

A low flag rate can still hide an unfair one

A low overall flag rate can hide a few claim cohorts that carry most of the flags a reviewer later overturns — and one cohort flagged far more often than the rest. Nestack reports performance by slice, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Overseas multi-currency trips5.4%3.4× Review
Card-and-cash claim pairs4.2%2.6× Review
New joiners' first claims3.0%1.9× Watch
Routine domestic claims1.1%0.7× Normal
Bar: overturn lift vs. domestic-claim baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

A wrong flag retires the rule behind it

Each cycle retires a rule that flagged the wrong people and keeps the claim that proved it, and the flag-rate comparison is re-run before it ships.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Overturn rates rise, or one team's flag rate pulls away from the rest.

02Diagnose

Traced to a rule, a policy version, a receipt read or a card-feed join.

03Improve

The rule or its threshold is re-approved with finance and version-linked.

04Verify

Re-run on held-out claims, then re-checked for flag rates by team.

05Learn

The overturned claim becomes a regression case; the reason is rewritten.

Learn → DetectThe return edge. Every rule that flagged the wrong person leaves a test behind.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, policy encoding, detection logic, evaluation and fairness testing, integration, then a shadow cycle and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Workflow discovery and automation-boundary definition.
02Expense-platform and card-feed assessment.
03Expense-policy encoding and rule library.
04Claim, receipt and card-feed ingestion.
05Duplicate, split-claim and per-diem detection.
06Mileage and VAT/GST reclaim checks.
07Confidence scoring and flag ranking.
08Flag explanations written for claimants.
09Reviewer workflow and contested flags.
10Evaluation suite, regression and fairness testing.
11Expense-platform and GL integration.
12Shadow cycle, deployment and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne policy, one entity ProductionProduction integration AdvancedMultiple entities / countries
Introduced at Pilot
Policy checks and flags
Human decision on every flag
Baseline evaluation
Introduced at Production
Duplicate and split-claim checks
Card-feed reconciliation
Claimant-facing flag reasons
Fairness reporting by slice
Observability and evaluation
Introduced at Advanced
Per-diem, mileage and VAT checks
Multi-country policy and currency
Enterprise controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on policy complexity, integrations, claim volume, the number of countries and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your expense policy and its past versions Expense-policy encoding, version history and the rule libraryWeek 1
02A representative sample of submitted claims and receipts Claim, receipt and card-feed ingestion, and the duplicate-detection baselineWeek 2
03Corporate-card feeds and access to the expense platform Expense-platform and card-feed assessment, then integration setupWeek 2
04Employee grade, entitlement and cost-centre data Grade-entitlement, per-diem and mileage rate checksWeek 3
05Who decides a flag, and who hears a contested one Automation-boundary definition, reviewer workflow and contested flagsWeek 4
06Claims you have previously upheld and previously waived Evaluation suite, regression cases and fairness testingWeek 4
07Named reviewers and a team willing to run a shadow cycle Flag explanations for claimants, then the shadow cycle and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Phases are drawn over the weeks they actually occupy. Week 5 runs the agent in shadow against a live expense cycle while evaluation is still open.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Workflow discovery, policy-version mapping and the automation boundary W2Claim, receipt and card-feed ingestion, and the detection baseline W3Detection logic, confidence scoring and flag explanations W4Evaluation suite, fairness testing and the reviewer workflow W5Expense-platform integration and a shadow run on live claims W6Shadow flags decided by named reviewers, verified, then handover
Reading the bandBars cover only the weeks their work is named in. Evaluation stays open into week 5 because the shadow run is what tests it.
At the end of W6One expense cycle has been flagged in shadow and every flag decided by a named reviewer, then Agent Care takes over monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Accounting AI agent

Build an expense-audit agent around your policy and your people.

Show us your expense policy, your card feeds and who decides a contested flag. We'll encode one version of the policy, test it against claims you have already settled and compare the flags.

Nestack Agents · Expense audit & policy complianceAGT-ACC-08 · Agent Care available after launch