Nestack Agent Care
Industries / Operations / Document processing agent

Operations AI agent · Document processing

Back-Office Document Processing AI Agent

Turn the inbound post into structured fields, score each field against the page region it was read from, keep the document as the evidence, and hand the record to a named reviewer.

4–6 weeksTypical delivery
Your stackDeployment
Original keptNamed reviewer
Agent CareAfter launch

What this agent does

Reads the document, never accepts the record

In
01

A document arrives by post, portal, email or scan, and it is registered before a field is read out of it.

02

A field is read, and the page and the region it came off travel with the value from there on.

Reason
03

A value comes back confident, and the score is filed as a control on the read, not as a finding that it is right.

04

A supplier name lifts cleanly off a letterhead belonging to the printer, and the field is wrong at full confidence.

05

A document says invoice and behaves like a statement, so the type it announces is treated as a claim as well.

Decide
06

A duplicate arrives on a second channel in another format, and it is matched back rather than read in again.

07

A handwriting sits over a printed line, and it is raised as unread rather than quietly passed over.

Out
08

A record goes to a reviewer, and the document it came from stays attached to it, page by page.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

Reading, scoring and holding belong to the agent. Accepting a field, posting the record and paying against it belong to a named reviewer.

Example workflow

One document, post to acceptance

AgentHuman
1Document receivedPostal scan, supplier email, portal download, fax or a shared-drive drop
2Type and pages settledWhat the document appears to be, where it starts and ends, and the channel it came in on
3Fields read and scoredThe value, the page and region it came off, the score and what could not be read
4Controls appliedDuplicate checks, type checks, in-context checks and field confidence
No human action required

Stages 1 to 4 run unaided, and nothing is posted at any of them — the agent is reading, and the reviewer lane opens at the confidence gate.

5DecisionSplits at the confidence gate
Above the threshold

Goes to the named reviewer to accept.

Anything below

Adds a second read first.

Reviewer acceptance

The record is held with its fields, their scores and the document they were read from.

Accept · Correct · Send to second read
Accepted — by the named reviewer
6Back-office records updatedOnly where write access and records policy allow it
7Outcome evaluatedField accuracy, score calibration, reviewer corrections and what review found
Corrections

Each reviewer correction is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Approving a document for payment.
Posting an extracted record to a system of record.
Deciding a low-confidence field is good enough.
Paying anything against a document.
Automation boundaryAgent acts unaided
Read whatever arrives into a structured record.
Score each field and carry the page region it was read from with it.
Keep the original beside the record it produced.
Route anything below the threshold to a second read, and hold the rest.
Nothing is posted or paid except by a named reviewer, inside the agreed boundaries.
Judging what a document actually is.
Accepting a field no page supports.
Setting the threshold fields are accepted at.
Changes to extraction rules or thresholds.

Example output

One inbound document, annotated

This serves a back office that has to answer, months later, where a figure in a record came from; below is one document exactly as the agent leaves it.

Extraction output · single documentIllustrative example
Document
Recorded as
Read from
Evidence of record
Confidence
Held for
Supplier invoice, scanned post
Supplier name, read confident
Page one, letterhead
Original scan, 4 June 2026
Held unaccepted
The named reviewer, by name
As receivedTaken from the scanned document on file, and it asserts nothing the page does not carry.
What the record holds Page and region Original scan Arrival channel
Why no acceptance hereDeciding a field is good enough is a judgement the reviewer makes.
ActionAcceptCorrectSend to second read
What the score decidesBelow the configured threshold the field gets a second read before the reviewer sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every documentFrom the channel it arrived on
03Evidence

Where the original is held

Freight document processing reads bills of lading, rate confirmations and customs paperwork; this is the back office inbound post, whatever it happens to be.

01Approved path

A field is a claim

A supplier name read cleanly off a letterhead can belong to the printer rather than the supplier: the read was good, the score was high, and the field is still wrong.

02Human review

What was checked, and not found

Checked across the document types in scope: nothing published says what a confidence score has to mean, and nothing sets the level a field may be accepted at, so the threshold is a product decision that gets written down and reported by document type rather than in total.

04Build an evidence trail

The document, the field extracted from it and the reviewer who accepted it stay together.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Inbound channelsPost room scans · shared mailboxes
Whatever arrives, in whatever shape
Capture and imagingKofax · ABBYY · scanner fleets
Page images and the text under them
Back-office systemsSAP · Oracle · NetSuite · Workday
Where an accepted record is written

Agent

Back-office document processing

Reads the document
Scores each field
Holds for the reviewer

Content and archiveSharePoint · OpenText · Box
Where the original document is kept
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six graders between the model and the record

Six graders in a row, the last the harshest. What is accepted is drawn in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeRoute the whole document to a human reviewer when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt and extraction rules, and note the version each field was read under.Track
L4TraceabilityRecord each document, the fields read from it, the page each came off and every reviewer action.Record
L3Reviewer releaseHold each field for a named reviewer; the hold governs acceptance, not whether the value is right.Gate
L2Field guardrailsTest each field against the page region under it and the document around it, and refuse a value impossible in that context.Restrict
L1Confidence thresholdsRoute a field below the configured threshold to a second read before the record reaches the reviewer.Require review
Model coreFields read — the value, the page it came off, the score and what could not be read
L1 – L2Test whether a field may stand
L3Leaves the acceptance to a named reviewer
L4 – L5Keep the field and the document behind it
L6Routes the whole document to a human when signals degrade

How Nestack evaluates it

Evaluate the whole read — not only the field that comes out of it.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the record the back office works from
Depth of coverage ▼
E1Final-output evaluationDid each field match the page region it was read from?
E2Step-level evaluationDid the agent read the right document, the right pages and the live extraction rules?
E3Tool evaluationDid it read and write the correct document and the correct field?
E4Confidence calibrationDo low-confidence fields actually attract more reviewer corrections?
E5Slice evaluationHow does performance change across specific document types?
E6Business outcomeHow many fields needed a correction before the reviewer accepted?
Floor — the document a field rests on

Failure modes

Where each failure originates in the agent

Seven failure modes, each placed at the stage where the read goes wrong.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
NL-03

Annotation never read

A handwritten change over a printed line is missed.

Stage gathersThe pages, the regions, the text and the type
02 · Reasoning2 modes
NL-04

Confident and wrong

A clean read returns the wrong field.

NL-06

Typed by its own letterhead

A statement is taken for the invoice it names.

Stage proposesThe fields, the scores, the type and the gaps
03 · Tool / write2 modes
NL-02

Thin field passed forward

A field moves on without the second read.

NL-05

Duplicate read as new

The same document arrives in two formats.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
NL-01

Accepted, page reference missing

The record shows a value but not where it came from.

Stage returnsThe fields a back office works from and a person accepts
05 · Change / Version1 mode
NL-07

Silent extraction drift

A model change widens what the agent will read off a page.

Stage tracksModel, prompt, field rules and thresholds
Sev-1 · a record posted with no reviewer Sev-2 · a wrong field reaches the record Sev-3 · a scan degrades, document held back

Affected slices

Handwritten forms absorb the corrections

A type-level review-rate figure can read clean while handwritten and annotated forms carry most of the corrections. Nestack reports the correction rate by document type, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Handwritten and annotated forms8.9%3.7× Review
Multi-page batch scans6.4%2.7× Review
Statements and credit notes4.0%1.7× Watch
Structured supplier invoices1.7%0.7× Normal
Bar: correction-rate lift vs. structured-invoice baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What a confident wrong field costs

A cycle closes when the field accepted below threshold is a regression case. That suite is what the next batch cleared is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Correction rate rises on handwritten and annotated forms.

02Diagnose

The invoice whose supplier name came back confident and wrong is taken apart page by page until the read that produced it is found.

03Improve

The change ships numbered, with the documents that caused it attached.

04Verify

One document case still failing is enough to stop the release.

05Learn

One case joins the suite, one line joins the extraction rules.

Learn → DetectThe return edge. The next batch cleared runs against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, field extraction, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Document-inventory and automation-boundary work.
02Inbound channel and capture sources.
03Field-definition and confidence-threshold mapping.
04Document ingestion and page splitting.
05Field, page and document binding.
06Confidence scoring and review routing.
07Reviewer acceptance workflow.
08Back-office system integration.
09Extraction and confidence cases.
10Guardrails and review controls.
11Document-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne document type, one site ProductionProduction intake workflow AdvancedMultiple types / channels
Introduced at Pilot
Field extraction to your record
Named reviewer acceptance
Document-type baseline
Introduced at Production
Reporting by document type
Reviewer workflow in your systems
Approved write-back
Inbound-channel integration
Introduced at Advanced
Multi-format duplicate matching
Cross-entity document packs
Large daily intake
Multi-language extraction controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, document volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your document types and the fields each one has to yield Field definition and page-region bindingWeek 1
02Representative documents from each inbound channel Ingestion, page splitting and the extraction baselineWeek 2
03Your acceptance thresholds and the reviewers they name Field mapping, threshold setting and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Inbound channel and capture assessment, then integration setupWeek 2
05Extractions you would not want re-keyed Confidence cases and failure-mode testingWeek 4
06What no extracted field may prove Confidence scoring, review routing, guardrails and acceptance controlsWeek 3
07A named reviewer who accepts the record Acceptance workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

No band was padded to reach the next column; the pair in week five is a real overlap, not a gap.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Document-type discovery, field definition and the automation boundary W2Channel integration and the extraction baseline W3Page splitting, confidence logic and acceptance controls W4Evaluation suite, confidence cases and failure-mode testing W5Back-office system integration, pilot intake and targeted corrections W6One month of intake run under the back-office lead, then Agent Care handover
Reading the bandA band sits on the weeks its own work is named for and no others, and week five is genuinely shared.
At the end of W6When the extraction record validates, Agent Care takes the agent on.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Operations AI agent

Build a document processing agent around the field that came back wrong with a green tick beside it.

Show us a week of inbound post and the fields your team re-keys out of it. We will read one batch as it arrived, mark where a confident field and the page underneath it disagree, and hand each field back to a named reviewer.

Nestack Agents · Document processingAGT-OP-07 · Agent Care available after launch