Nestack Agent Care
Industries / Accounting / Bookkeeping agent

Accounting AI agent · Bookkeeping

Bookkeeping & Transaction-Classification AI Agent

Classify transactions, suggest ledger accounts, handle ambiguous entries and route exceptions for review — with evaluation, audit trails, guardrails and human approval built in.

4–6 weeksTypical delivery
Your stackDeployment
Exception-basedHuman review
Agent CareAfter launch

What this agent does

Automates the repetitive classification layer

In
01

Ingest transaction records from supported accounting, banking, CSV or API sources.

02

Normalise transaction descriptions and metadata.

Reason
03

Identify likely transaction type and suggest the ledger account or category.

04

Apply client-specific chart-of-account mappings and classification rules.

05

Use approved supplier history, policies and prior classifications as context.

Decide
06

Detect ambiguous, unusual or policy-sensitive transactions.

07

Route low-confidence entries to human review.

Out
08

Retain classification rationale, source context and reviewer corrections.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

The agent proposes or executes actions according to the approval boundaries agreed during implementation.

Example workflow

One transaction, end to end

AgentHuman
1Transaction receivedBank feed, CSV, accounting system or transaction API
2Context gatheredMerchant, description, amount, client history, chart of accounts and approved policy
3Classification proposedAccount, category, tax treatment and confidence
4Controls appliedPolicy checks, duplicate checks, unusual-value checks and confidence threshold
No human action required

Stages 1 to 4 run without a person in the loop — review is exception-based, so the lane stays empty until the confidence gate.

5DecisionSplits on the confidence threshold
High confidence

Follows the approved path.

Low confidence

Enters human review.

Human review

The entry is held with its rationale, source context and confidence.

Approve · Correct · Request review
Approved — handed back
6Accounting system updatedOnly when write access and approval policy allow it
7Outcome evaluatedClassification correctness, overrides, exceptions and business outcome
Overrides

Corrections made in review are counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Unusual high-value entries.
Transactions with insufficient context.
New suppliers or new categories.
Intercompany transactions.
Automation boundaryAgent acts unaided
Classify routine transactions and suggest the ledger account.
Apply approved chart-of-account mappings and rules.
Use approved supplier history and prior classifications.
Route low-confidence entries to review and retain the rationale.
Write actions run only inside the approval boundaries agreed during implementation.
Manual journals or corrections with material impact.
Uncertain tax treatment.
Policy exceptions.
Changes to chart-of-account rules or automation thresholds.

Example output

One line in the ledger, annotated

Everything the agent proposes is attached to the record it came from.

Classification output · single recordIllustrative example
Merchant
Description
Amount
Suggested account
Confidence
Tax treatment
Cloud software vendor
Annual software subscription
$4,800.00
Software subscriptions
94%
Standard configured treatment
As receivedTaken from the transaction source and normalised — nothing on this side is inferred.
Evidence used Approved supplier match Prior mapped transactions Description pattern
Why this treatmentFrom the client's configured treatment for that account — not decided per transaction.
ActionApproveCorrectRequest review
What the score decidesBelow the configured threshold the entry routes to human review instead of the approved path.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
All transactionsFrom the transaction source
03Classification

Apply client-specific context

Use approved chart-of-account mappings, supplier history and accounting policies.

01Approved path

Reduce routine coding workload

Routine transaction classification is handled automatically or semi-automatically.

02Human review

Focus accountants on exceptions

Low-confidence and unusual entries move to review instead of every transaction requiring manual attention.

04Build an evidence trail

Retain source context, recommendation, confidence, evaluator result and human correction — on both paths.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Accounting systemsQuickBooks · Xero · Sage
NetSuite · ERP/accounting APIs
Banking / transaction dataBank feeds · CSV
Transaction APIs
DocumentsInvoices · receipts
Supporting documents

Agent

Bookkeeping & transaction classification

Reads context
Proposes classification
Routes exceptions

WorkflowEmail · ticketing
Approval systems
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and your ledger

Each control wraps the one inside it. A classification clears every layer before anything is written.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeRestrict automation if evaluations or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, rules and agent configuration changes.Track
L4TraceabilityRecord source, recommendation, evaluation, action, and human override.Record
L3Human approvalDefine which transactions can be processed automatically.Gate
L2Policy guardrailsRestrict classifications to approved accounts and rules.Restrict
L1Confidence thresholdsLow-confidence classifications require review.Require review
Model coreClassification proposed — account, category, tax treatment and confidence
L1 – L2Decide whether the answer may stand
L3Decides whether a human must sign it
L4 – L5Keep the record and the version honest
L6Pulls automation back when signals degrade

How Nestack evaluates it

Evaluate the full workflow — not only the final classification.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the answer the client sees
Depth of coverage ▼
E1Final-output evaluationWas the account/category classification correct?
E2Step-level evaluationDid the agent use the correct context, mapping and accounting policy?
E3Tool evaluationDid it read or write to the correct system and record?
E4Confidence calibrationDo low-confidence decisions actually contain more errors?
E5Slice evaluationHow does performance change across specific transaction groups?
E6Business outcomeHow many transactions required correction or manual intervention?
Floor — the outcome the client pays for

Failure modes

Where each failure originates in the agent

Seven failure modes plotted against the five stages of the agent lifecycle.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
BK-03

Context retrieval failure

Relevant supplier/client history is missing.

Stage gathersMerchant, description, client history and chart of accounts
02 · Reasoning2 modes
BK-04

Tax-treatment error

Transaction receives incorrect configured tax handling.

BK-06

Policy-rule violation

Suggested treatment conflicts with client policy.

Stage proposesAccount, category, tax treatment and confidence
03 · Tool / write2 modes
BK-02

Low-confidence auto-posting

Agent executes despite insufficient certainty.

BK-05

Duplicate processing

Same transaction is classified or posted twice.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
BK-01

Transaction misclassification

Wrong ledger/category chosen.

Stage returnsThe classification the client and the ledger see
05 · Change / Version1 mode
BK-07

Silent version regression

Model or configuration change reduces classification quality.

Stage tracksModel, prompt, rules and configuration changes
Sev-1 · acts outside the boundary Sev-2 · wrong entry reaches the ledger Sev-3 · input degrades, entry routes to review

Affected slices

Overall health can hide concentrated risk

Aggregate classification quality can look acceptable while a small number of transaction cohorts carry most of the failures. Nestack reports performance by slice, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
High-volume bank feeds4.6%3.1× Review
Ambiguous expense categories3.6%2.4× Review
New-client first close2.8%1.9× Watch
Established ledgers0.9%0.6× Normal
Bar: lift vs. established-ledger baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

The loop does not end at Learn

Each completed cycle leaves the agent with one more regression case and one more playbook entry, which is what the next detection is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Misclassification rate increases in a slice.

02Diagnose

Failure isolated to context, rule, prompt or workflow.

03Improve

Approved corrective action is recorded, version-linked and implemented.

04Verify

Affected regression cases are re-run.

05Learn

Failure becomes a permanent test and playbook update.

Learn → DetectThe return edge. The next cycle starts against a larger regression suite than the last one.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, data, agent workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Workflow discovery and automation-boundary definition.
02Source-system and API assessment.
03Chart-of-account and classification-policy mapping.
04Transaction ingestion and normalisation.
05Context retrieval and classification logic.
06Confidence scoring and exception routing.
07Human review workflow.
08Accounting-system integration.
09Evaluation suite and regression cases.
10Guardrails and approval controls.
11Observability and trace instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne transaction source ProductionProduction integration AdvancedMultiple entities / systems
Introduced at Pilot
Classification recommendations
Human approval
Baseline evaluation
Introduced at Production
Client-specific rules
Review workflow
Approved write actions
Observability and evaluation
Introduced at Advanced
Complex policies
Multi-stage approvals
High volume
Enterprise controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Chart of accounts and relevant category mappings Chart-of-account and classification-policy mappingWeek 1
02Representative historical transactions Transaction ingestion and normalisation, and the classification baselineWeek 2
03Approved accounting policies and exception rules Accounting-policy mapping and automation-boundary definitionWeek 1
04Access to relevant APIs, feeds or exports Source-system and API assessment, then data and integration setupWeek 2
05Examples of difficult or ambiguous classifications Evaluation suite, regression cases and failure-mode testingWeek 4
06Approval thresholds and write-action boundaries Confidence scoring, exception routing, guardrails and approval controlsWeek 3
07Named reviewers or test users Human review workflow, then pilot workflow and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Phases are drawn over the weeks they actually occupy. Week 5 carries both evaluation and launch work.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Workflow discovery, accounting-policy mapping and automation boundary W2Data/integration setup and classification baseline W3Agent workflow, confidence logic and human-review controls W4Evaluation suite, guardrails and failure-mode testing W5Integration, pilot workflow and targeted corrections W6Production validation, verification and Agent Care handover
Reading the bandBars span only the weeks their work is named in. The week 5 overlap is real, not padding.
At the end of W6Production validation and verification complete, then Agent Care takes over monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Accounting AI agent

Build a bookkeeping agent around your accounting workflow.

Show us your transaction sources, chart of accounts and review process. We'll map the workflow, identify the automation boundary and recommend the safest path to production.

Nestack Agents · Bookkeeping & transaction classificationAGT-ACC-01 · Agent Care available after launch