Run the tests of details you scope — population checks, documented selections, three-way matching and journal-entry work — and write up every exception, with evaluation, audit trails and the sufficiency judgement left with the auditor.
Ingest the population extract, the audit programme step and the documents the test calls for.
02
Normalise each item into entity, period, account, assertion and the test it belongs to.
Reason
03
Agree the extract to the ledger control total before a single item is drawn from it.
04
Draw the selection the auditor specified, and record the method, parameters and seed used.
05
Perform the defined test on every item selected — match, recalculate, agree the date, reperform.
Decide
06
Flag exceptions, items the test could not be applied to, and evidence that misses the assertion.
07
Return every exception and every untested item to the auditor instead of clearing it.
Out
08
Retain the extract, the selection, each item tested, the evidence seen and the result.
09
Write to the workpaper file only inside the approval boundaries agreed during implementation.
→Product statement
The agent executes the procedures the auditor scoped and documents what it found; it gives no assurance, and the auditor judges sufficiency and forms every conclusion.
Example workflow
One procedure, end to end
AgentHuman
1Procedure receivedAudit programme step, workpaper request or the engagement team's own scope
2Scope resolvedEntity, period, account, assertion, population definition and the selection the auditor set
3Test performedExtract agreed to the control total, the specified selection drawn, then each item tested
4Controls appliedPopulation-completeness check, evidence-to-assertion check, source-data checks and confidence threshold
No human action required
Stages 1 to 4 run without a person in the loop — nothing has been relied on yet, so the lane stays empty until the gate.
5DecisionSplits on the confidence and exception gate
High confidence
Recorded as a worked step with its evidence.
Low confidence
Held with the exception it could not resolve.
Auditor review
The step is held with its population, selection, evidence and the exception that stopped it.
Accept · Amend · Extend testing
Accepted — handed back▼
6Step written to the fileRecorded as work performed, never as a conclusion; the auditor accepts the step
7Outcome evaluatedException precision, population-completeness checks, auditor overturns and business outcome
Overturns
Amendments and extended testing at review are counted in the evaluation.
What should not run autonomously
Human approval stays in control
Outside the boundary — human approval required8 items
Whether the evidence is sufficient and appropriate.
Sample size, method and materiality decisions.
Any conclusion on an assertion.
Whether an exception is a misstatement.
Automation boundaryAgent acts unaided
✓Agree the population extract to the ledger control total.
✓Draw the selection the auditor specified and record how.
✓Perform the defined test on every item selected.
✓Document each exception and hand it to the auditor.
Workpapers are written only inside the approval boundaries agreed during implementation.
Treating an exception as an isolated anomaly.
Accepting entity-produced data as reliable.
Extending, reducing or stopping a test.
Sign-off of any workpaper or audit step.
Example output
One tested item, annotated
Everything the agent records is attached to the procedure and the population it came from.
Worked step · single procedureIllustrative example
Procedure
Population
Period
Result
Confidence
Population check
Three-way match
Posted supplier invoices
FY25
Two open exceptions
89%
Agrees to the control total
As receivedThe procedure, population and period as the audit programme defines them — the agent redefines neither.
Journal-entry work and scanned evidence behave nothing like a three-way match, and they carry most of the missed and mis-flagged exceptions. Nestack reports performance by slice, not only in total.
Slice performance — reported separately, not only in aggregateIllustrative example
Slice
Failure rate
Lift
Lift vs. threshold
Status
Journal-entry testing runs
6.3%
3.5×
Review
Scanned-document evidence
5.4%
3.0×
Review
Confirmation exception runs
3.4%
1.9×
Watch
Three-way match runs
0.9%
0.5×
Normal
Bar: lift vs. three-way-match baseline · scale 0–4.0× · tick marks the 2.0× review threshold2 of 4 slices over threshold
Evidence-linked improvement
Scoring comes from the auditor, not the agent
What the reviewer changed — a missed exception, or one raised that was not real — is what the next release is measured against.
Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect
Overturns rise in a test cohort, or exception precision drops.
02Diagnose
Failure isolated to the extract, the draw, the evidence match or the rule.
03Improve
The rule or the match logic is corrected — approved, version-linked and dated.
04Verify
Archived engagement data is re-run and the two results are compared.
05Learn
The overturned item becomes a case in the regression suite.
Learn → DetectThe return edge. A release that cannot reproduce last cycle's overturns does not ship.
Typical build scope
Twelve workstreams across six weeks
The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, populations and extracts, test execution, evaluation, integration, then production validation and handover.
WorkstreamWeek 1Week 2Week 3Week 4Week 5Week 6
01Procedure scoping and assurance-boundary definition.
02Source-system and workpaper assessment.
03Population definition and control-total mapping.
04Extract reconciliation and selection logic.
05Test execution, matching and recalculation.
06Exception rules and confidence scoring.
07Auditor review and acceptance workflow.
08Workpaper and document-system integration.
09Evaluation suite and regression cases.
10Guardrails and documentation controls.
11Observability and trace instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallelFinal scope and sequence confirmed in discovery
Engagement tiers
What each tier includes
Rows are the capabilities named in each tier's scope. Higher tiers include everything below them. No tier adds assurance or the sufficiency judgement.
Capability✓ in scope · — not at this tierPilotOne procedureProductionProduction integrationAdvancedGroup engagements
Introduced at Pilot
Defined-procedure execution✓✓✓
Population completeness checks✓✓✓
Auditor acceptance✓✓✓
Baseline evaluation✓✓✓
Introduced at Production
Reproducible selection records—✓✓
Exception write-ups—✓✓
Approved workpaper writes—✓✓
Observability and evaluation—✓✓
Introduced at Advanced
Group and multi-entity testing——✓
Multi-stage review——✓
Enterprise controls——✓
Build priceFrom $5,000From $8,000Custom quote
Final build priceConfirmed after discovery based on the procedures in scope, population volumes, source systems, review controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.
What we need from you
What you bring, and what we build with it
Each input maps to a piece of build scope and a week in the delivery timeline.
You bringWe build with it
01The procedures in scope and the programme steps they sit under→Procedure scoping and the population definition for eachWeek 1
02The assurance boundary and who accepts each step→Assurance-boundary definition and acceptance controlsWeek 1
03Access to the ledger, the extracts and the document store→Source-system and workpaper assessmentWeek 2
04Population definitions, control totals and materiality→Extract reconciliation and control-total mappingWeek 2
05Selection methods and the exception thresholds you use→Selection logic, exception rules and confidence scoringWeek 3
06Exceptions your reviewers overturned last year→Evaluation suite, regression cases and failure-mode testingWeek 4
07Named auditors or test users→Auditor acceptance workflow, then pilot testing and production validationWeeks 5–6
Nothing else is requiredDeployment, documentation and Agent Care handover are ours.
Delivery timeline
Four phases across six weeks
Phases are drawn over the weeks they actually occupy. Weeks 5 and 6 run a real audit step, which is the only place the exception rules meet live data.
PhaseW1W2W3W4W5W6
DiscoveryW1
BuildW2 – W3
EvaluateW4 – W5
Pilot & LaunchW5 – W6
Week focusW1Procedure scoping, assurance boundary and population definitionsW2Ledger and document access, extract reconciliation and the selection baselineW3Test execution, exception rules and auditor acceptance controlsW4Evaluation suite, documentation guardrails and failure-mode testingW5Workpaper integration, a pilot procedure and targeted correctionsW6One audit step run end to end, verification and Agent Care handover
Reading the bandA bar covers only the weeks its work is named in. Week 5 carries two because the pilot procedure runs against the evaluation suite, not after it.
At the end of W6One procedure has been run end to end and accepted by a named auditor, then Agent Care monitors population checks, exception precision and overturns.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.
Next step · Accounting AI agent
Build an evidence-testing agent around your audit programme.
Show us the procedures you want run, the populations they draw from and who accepts each step. We'll take one procedure end to end, agree what the agent may never conclude and scope the rest from there.