Classify transactions, suggest ledger accounts, handle ambiguous entries and route exceptions for review — with evaluation, audit trails, guardrails and human approval built in.
Aggregate classification quality can look acceptable while a small number of transaction cohorts carry most of the failures. Nestack reports performance by slice, not only in total.
Slice performance — reported separately, not only in aggregateIllustrative example
Slice
Failure rate
Lift
Lift vs. threshold
Status
High-volume bank feeds
4.6%
3.1×
Review
Ambiguous expense categories
3.6%
2.4×
Review
New-client first close
2.8%
1.9×
Watch
Established ledgers
0.9%
0.6×
Normal
Bar: lift vs. established-ledger baseline · scale 0–4.0× · tick marks the 2.0× review threshold2 of 4 slices over threshold
Evidence-linked improvement
The loop does not end at Learn
Each completed cycle leaves the agent with one more regression case and one more playbook entry, which is what the next detection is measured against.
Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect
Misclassification rate increases in a slice.
02Diagnose
Failure isolated to context, rule, prompt or workflow.
03Improve
Approved corrective action is recorded, version-linked and implemented.
04Verify
Affected regression cases are re-run.
05Learn
Failure becomes a permanent test and playbook update.
Learn → DetectThe return edge. The next cycle starts against a larger regression suite than the last one.
Typical build scope
Twelve workstreams across six weeks
The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, data, agent workflow, evaluation, integration, then production validation and handover.
WorkstreamWeek 1Week 2Week 3Week 4Week 5Week 6
01Workflow discovery and automation-boundary definition.
02Source-system and API assessment.
03Chart-of-account and classification-policy mapping.
04Transaction ingestion and normalisation.
05Context retrieval and classification logic.
06Confidence scoring and exception routing.
07Human review workflow.
08Accounting-system integration.
09Evaluation suite and regression cases.
10Guardrails and approval controls.
11Observability and trace instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallelFinal scope and sequence confirmed in discovery
Engagement tiers
What each tier includes
Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.
Capability✓ in scope · — not at this tierPilotOne transaction sourceProductionProduction integrationAdvancedMultiple entities / systems
Introduced at Pilot
Classification recommendations✓✓✓
Human approval✓✓✓
Baseline evaluation✓✓✓
Introduced at Production
Client-specific rules—✓✓
Review workflow—✓✓
Approved write actions—✓✓
Observability and evaluation—✓✓
Introduced at Advanced
Complex policies——✓
Multi-stage approvals——✓
High volume——✓
Enterprise controls——✓
Build priceFrom $5,000From $8,000Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.
What we need from you
What you bring, and what we build with it
Each input maps to a piece of build scope and a week in the delivery timeline.
You bringWe build with it
01Chart of accounts and relevant category mappings→Chart-of-account and classification-policy mappingWeek 1
02Representative historical transactions→Transaction ingestion and normalisation, and the classification baselineWeek 2
03Approved accounting policies and exception rules→Accounting-policy mapping and automation-boundary definitionWeek 1
04Access to relevant APIs, feeds or exports→Source-system and API assessment, then data and integration setupWeek 2
05Examples of difficult or ambiguous classifications→Evaluation suite, regression cases and failure-mode testingWeek 4
06Approval thresholds and write-action boundaries→Confidence scoring, exception routing, guardrails and approval controlsWeek 3
07Named reviewers or test users→Human review workflow, then pilot workflow and production validationWeeks 5–6
Nothing else is requiredDeployment, documentation and Agent Care handover are ours.
Delivery timeline
Four phases across six weeks
Phases are drawn over the weeks they actually occupy. Week 5 carries both evaluation and launch work.
PhaseW1W2W3W4W5W6
DiscoveryW1
BuildW2 – W3
EvaluateW4 – W5
Pilot & LaunchW5 – W6
Week focusW1Workflow discovery, accounting-policy mapping and automation boundaryW2Data/integration setup and classification baselineW3Agent workflow, confidence logic and human-review controlsW4Evaluation suite, guardrails and failure-mode testingW5Integration, pilot workflow and targeted correctionsW6Production validation, verification and Agent Care handover
Reading the bandBars span only the weeks their work is named in. The week 5 overlap is real, not padding.
At the end of W6Production validation and verification complete, then Agent Care takes over monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.
Next step · Accounting AI agent
Build a bookkeeping agent around your accounting workflow.
Show us your transaction sources, chart of accounts and review process. We'll map the workflow, identify the automation boundary and recommend the safest path to production.