Nestack Agent Care
Industries / Procurement / Purchasing / Spend analysis agent

Procurement AI agent · Spend analysis

Spend Analysis AI Agent

Place each transaction in a category with the rule that placed it visible, consolidate the supplier records that are one supplier, name what would not classify, and reconcile the total to the ledger.

4–6 weeksTypical delivery
Your stackDeployment
Rule visibleAnalytics owner
Agent CareAfter launch

What this agent does

Places the transaction, never sets the taxonomy

In
01

A transaction lands from the ledger or a card feed, and the rule that will place it is written down beside it.

02

A category is decided, not described — the placement settles who owns the spend and whether it goes to market.

Reason
03

A line codes to miscellaneous, and it is reported as unplaced rather than folded into a neighbouring category.

04

A supplier appears under several names, and the records that are one relationship are proposed for consolidation.

05

A merge rests on a name match alone, and it is held for the analytics owner with the evidence for it shown.

Decide
06

A cut is drawn and reconciled to the ledger, because a cut that will not reconcile is disbelieved when challenged.

07

A taxonomy stops matching how the business buys, and the drift is raised rather than absorbed by the mapping.

Out
08

A trend crosses a taxonomy change, and the change is marked on the series instead of being smoothed through.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

Classification, consolidation and reconciliation belong to the agent. The taxonomy, the ownership of a category and any claim of a saving belong to a named analytics owner.

Example workflow

One transaction, ledger to cut

AgentHuman
1Transaction receivedGeneral ledger, payables, card feeds, expense claims or supplier invoices
2Context assembledSupplier name and identifiers, line description, cost centre, ledger account and prior placements
3Category proposedCategory, supplier record, the rule that placed it and confidence
4Controls appliedTaxonomy checks, supplier-match checks, ledger reconciliation checks and confidence threshold
No human action required

Stages 1 to 4 run unaided, and no cut is published at any of them — the agent is placing, and the analytics owner lane opens at the confidence gate.

5DecisionSplits at the confidence gate
Placed on a matched rule

Goes to the analytics owner to accept.

Unplaced or a weak supplier match

Adds a category manager read first.

Analytics owner review

The transaction is held with its rule, the supplier records behind it and the confidence.

Accept · Recode · Send to category review
Accepted — by the analytics owner
6Spend and reporting records updatedOnly where write access and records policy allow it
7Outcome evaluatedPlacement correctness, unplaced spend, merge corrections and what review found
Corrections

Each recode made in review is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Setting the taxonomy a company reports its spend at.
Deciding who owns a category and its budget.
Committing a supplier merge without a named review.
Declaring a saving against a prior period.
Automation boundaryAgent acts unaided
Place each transaction and show the rule that placed it.
Mark the transaction it could not place as unplaced.
Propose the supplier records that are one supplier, with the evidence.
Reconcile the classified total back to the ledger that produced it.
Nothing enters a published cut except by a named analytics owner, inside the agreed boundaries.
Judging whether a category is worth going to market.
Telling a board that a spend cut is complete.
Setting the threshold a category is reviewed at.
Changes to the taxonomy, the rules or the thresholds.

Example output

One transaction, annotated

This serves a procurement team who may have to defend a category cut to the people who own the budgets; below is one transaction exactly as the agent leaves it.

Classification output · single transactionIllustrative example
Supplier as billed
Line description
Ledger account
Category proposed
Confidence
Held for
Regional trading name
Advisory work on a monthly retainer
Miscellaneous other costs
Professional services, proposed
Held unaccepted
The analytics owner, by name
As receivedTaken from the ledger and the card feed as they stand — nothing on this side is rewritten by the agent.
Evidence used Ledger line Supplier name variants Prior placements
Why this categoryPlacing a line here decides who owns it and whether it goes to market.
ActionAcceptRecodeSend to category review
What the score decidesBelow the configured threshold a line picks up a category read before the owner sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every transactionFrom the ledger that carries it
03Classification

Where the spend is placed

Campaign spend pacing is a separate agent; this one reads third-party spend across the whole company, and its subject is the classification rather than the burn rate.

01Approved path

A category is a decision

Call a line professional services rather than contingent labour and you have decided who owns it and whether it ever goes to market.

02Human review

What was checked, and not found

No ledger consulted was built to a category structure — it answers to the finance close — so the mapping is reported with what it loses in each direction rather than as a clean fit.

04Build an evidence trail

The transaction, the category it was placed in and the rule that placed it stay together.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Ledger and payablesSAP · Oracle · NetSuite · Workday
General ledger and AP lines
Card and expense dataCorporate card feeds · claims
Expense, travel and mileage
Supplier and contract dataSupplier master · contract register
Registered names and identifiers

Agent

Spend analysis and classification

Reads the ledger
Places the transaction
Holds for the owner

Purchasing and invoicesCoupa · Ariba · Ivalua · Jaggaer
Purchase orders and invoice lines
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six sorters between the model and the cut

Six sorters in a line, the last the strictest. What is placed is drawn in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeNarrow the agent to listing unplaced transactions when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt and taxonomy rules, and note the version each placement was made under.Track
L4TraceabilityRecord each transaction, the rule that placed it, the supplier records merged and each read.Record
L3Owner releaseHold the cut for a named analytics owner; the hold governs release, not whether a placement is right.Gate
L2Taxonomy guardrailsTest each placement against the taxonomy in force, and return a line whose rule no longer stands.Restrict
L1Confidence thresholdsRoute a weak supplier match to a category manager read before the merge reaches a cut.Require review
Model coreTransaction placed — the category, the rule, the supplier record and confidence
L1 – L2Test whether a placement may stand
L3Leaves the acceptance to a named owner
L4 – L5Keep the transaction and the rule behind it
L6Leaves the transaction unclassified when signals degrade

How Nestack evaluates it

Evaluate the whole classification — not only the cut that comes out.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the cut a category review reads
Depth of coverage ▼
E1Final-output evaluationDid the placement record the rule it was actually made under?
E2Step-level evaluationDid the agent read the right ledger, the right period and the live taxonomy?
E3Tool evaluationDid it read and write the correct transaction and the correct category?
E4Confidence calibrationDo low-confidence placements actually attract more owner recodes?
E5Slice evaluationHow does performance change across specific categories?
E6Business outcomeHow many transactions needed a recode before the owner accepted?
Floor — the spend a budget owner answers for

Failure modes

Where each failure originates in the agent

Seven failure modes, each placed at the stage it first becomes visible.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
OK-03

Stale ledger extract read

The extract read is not the one now posted.

Stage gathersThe ledger, the cards, the periods and the names
02 · Reasoning2 modes
OK-04

Line placed on no rule

A category is assigned with no rule behind it.

OK-06

Retired category read as live

A withdrawn category is worked as the current one.

Stage proposesThe category, the rule and the supplier record
03 · Tool / write2 modes
OK-02

Weak merge passed forward

Two suppliers are joined on a name alone.

OK-05

Line counted twice in a cut

One transaction lands in two categories.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
OK-01

Placed, rule unrecorded

The cut shows a category but not the rule under it.

Stage returnsThe cut a category review and a board read
05 · Change / Version1 mode
OK-07

Silent taxonomy drift

A category widens while the stored cut keeps the old one.

Stage tracksModel, prompt, taxonomy rules and merge rules
Sev-1 · a cut published on no rule Sev-2 · a wrong category reaches the board Sev-3 · source degrades, line held unplaced

Affected slices

Miscellaneous absorbs the recodes

A category-level placement figure can read clean while miscellaneous-coded lines carry most of the recoding. Nestack reports the recode rate by slice, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Miscellaneous-coded lines8.2%3.7× Review
Fragmented supplier records5.9%2.7× Review
Services and contingent labour3.6%1.6× Watch
Established direct categories2.0%0.9× Normal
Bar: recode-rate lift vs. established-category baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What an unplaced line costs

The cycle ends when the transaction placed in the wrong category is a standing case. That suite is what the next cut published is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Recode rate rises on miscellaneous-coded lines.

02Diagnose

The invoice line coded to miscellaneous, on the one supplier the category manager had never heard of, is worked backwards until one cause is left standing.

03Improve

A change goes out numbered, and the cuts that forced it travel with it.

04Verify

Nothing ships while one classification case remains red.

05Learn

The case is kept, and the taxonomy rules are amended in that same commit.

Learn → DetectThe return edge. The next cut is measured against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, classification logic, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Taxonomy discovery and automation-boundary definition.
02Ledger, card and payables source review.
03Ledger-to-category and supplier-consolidation mapping.
04Transaction ingestion and normalisation.
05Placement logic and rule binding.
06Confidence scoring and review routing.
07Analytics owner review workflow.
08Ledger-and-card integration.
09Classification and taxonomy cases.
10Guardrails and placement controls.
11Spend-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne entity, one cycle ProductionProduction reporting workflow AdvancedMultiple entities / systems
Introduced at Pilot
Classification to your taxonomy
Analytics owner acceptance
Transaction-history baseline
Introduced at Production
Reporting by category
Category review workflow in your systems
Approved write-back
Ledger-and-card integration
Introduced at Advanced
Cross-entity supplier consolidation
Multi-source spend cubes
High transaction volume
Multi-entity taxonomy controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your live taxonomy and the categories it names Taxonomy capture and rule versioningWeek 1
02Representative ledger, card and invoice history Source binding, placement logic and the classification baselineWeek 2
03Your reporting calendar and the owners it names Taxonomy mapping, rule binding and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Ledger, card and payables assessment, then integration setupWeek 2
05Cuts you would not want re-derived Taxonomy cases and failure-mode testingWeek 4
06What no spend cut may prove Confidence scoring, review routing, guardrails and release controlsWeek 3
07A named analytics owner who accepts the cut Release to the analytics owner, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Each band is as wide as its phase costs, so week five carries a pair rather than a blank column.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Taxonomy discovery, rule mapping and the automation boundary W2Ledger and card integration and the classification baseline W3Placement logic, confidence scoring and release controls W4Evaluation suite, taxonomy cases and failure-mode testing W5Reporting integration, pilot cuts and targeted corrections W6One reporting cycle run under the analytics owner, then Agent Care handover
Reading the bandEach bar covers only the weeks its own work is named for, and week five is shared by design.
At the end of W6When the classification record validates, Agent Care assumes the agent.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Procurement AI agent

Build a spend analysis agent around the categories your last cut could not place.

Show us one spend cut you publish and the ledger behind it. Ask where the unmanaged suppliers live and the answer is miscellaneous, under a name nobody in the category team knew. We will map the taxonomy, set the automation boundary and name what stays with the analytics owner.

Nestack Agents · Spend analysisAGT-PR-05 · Agent Care available after launch