Nestack Agent Care
Industries / Retail & E-commerce / Order-status agent

Retail AI agent · Customer service

Customer-Service & Order-Status AI Agent

Answer order-status and policy questions from the live order record, hand anything legally operative to a named person, and never re-promise a date the merchant's own systems have not confirmed.

4–6 weeksTypical delivery
Your stackDeployment
Answers onlyEscalation lead
Agent CareAfter launch

What this agent does

Answers the question, never the promise

In
01

Pull the order, its per-line fulfilment state and the live carrier scan from supported commerce.

02

Timestamp every field the agent reads, and carry that as-of time forward into the answer.

Reason
03

Compose the reply from those fields, with the AI disclosure stated before anything else.

04

Apply the merchant's published shipping, return, warranty and privacy text as written.

05

Tie each stated fact to the field it was read from, and mark what the record does not say.

Decide
06

Detect a shipping promise, a warranty claim, a cancellation or a dispute inside the contact.

07

Route those contacts to the Customer Experience Escalation Lead instead of answering them.

Out
08

Retain the fields read, the reply, the disclosure and the routing against the ticket.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

The agent answers from the record; the escalation lead decides anything discretionary, and a delay notice or a refund stays with the people the merchant has named to author them.

Example workflow

One contact, arrival to answer

AgentHuman
1Contact receivedChat, email, voice or the helpdesk queue
2Context assembledFulfilment state per line, carrier scan, tender, policy version and subscription status
3Reply composedThe answer, its as-of time, any escalation trigger hit and confidence
4Controls appliedAI disclosure, source-field checks, escalation-trigger checks and confidence threshold
No human action required

Stages 1 to 4 run unaided and nothing has been sent — the lead's lane stays empty until a trigger fires or the confidence gate opens it.

5DecisionBranches at the confidence threshold
High confidence

Sends on the approved path.

Low confidence

Goes to the escalation lead first.

Escalation lead

The reply is held with the fields it read, its triggers and the confidence.

Send · Rewrite · Escalate to compliance
Released — answer sent
6Helpdesk record updatedOnly where write access and approval policy allow it
7Outcome evaluatedStatement accuracy against the record, escalation outcomes, reopens and post-answer corrections
Amendments

Every rewrite made in review is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Issuing a delay notice, or the renewed notice after it.
Authorising a refund, or stating when one will land.
Deciding a warranty claim or what its coverage reaches.
Denying, delaying or conditioning a subscription cancellation.
Automation boundaryAgent acts unaided
Report tracking from the live carrier field with its as-of time for the named owner.
Restate the published return.
Resend an order confirmation or invoice to the email.
Answer product-attribute questions from the live catalogue record.
Any write happens inside the boundaries agreed at implementation, never ahead of the lead's release.
Changing a shipping address, recipient or method after labelling.
Any statement about a bank or card dispute's timing or outcome.
Age-restricted, prescription, firearms or recalled-product contacts.
Granting a policy exception, goodwill credit or price adjustment.

Example output

One answer, annotated

Everything the agent states is attached to the field it was read from.

Order-status output · single contactIllustrative example
Order
Stated answer
Carrier status
Source of record
Confidence
Disclosure
Two-line order
One line is scanned and moving; the second is still unfulfilled
In transit
Live carrier feed
93%
Stated as AI at the start
As receivedTaken from the order record and the carrier feed at the time of reading.
Fields read Per-line fulfilment state Latest carrier scan Published shipping text
Why this answerIt reports what the feed said as of the read, and stops short of a delivery date.
ActionSendRewriteEscalate to compliance
What the score decidesBelow the configured threshold the reply routes to the escalation lead instead of.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every contactFrom the order record
03Answering

Answer from the record

Draw on the live order, carrier and published-policy fields the merchant already maintains.

01Approved path

Answer without the queue

Routine status and policy questions come back without a wait.

02Human review

Send the desk what carries risk

Promise, warranty, cancellation and dispute contacts are flagged, so the lead's time lands where exposure sits.

04Build an evidence trail

The question, the record it drew on and the routing stay with the ticket.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Commerce platform and OMSShopify
Salesforce Commerce Cloud
Post-purchase trackingNarvar
AfterShip · carrier APIs
Voice and CCaaSAmazon Connect
Genesys Cloud CX

Agent

Customer service & order status

Reads the order
Composes the answer
Routes escalations

Desk and subscriptionsZendesk · Gorgias
Recharge · Chargebee
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the customer

Each control wraps the one inside it. What a layer does not catch is named in the failure-mode map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeDrop the agent back to routing only when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, policy-text and escalation-rule changes.Track
L4TraceabilityRecord the fields read, the reply, the disclosure and the routing.Record
L3Escalation leadHold flagged replies for the named lead; it governs release, not whether a sent answer was right.Gate
L2Policy guardrailsTest replies against the published policy text and the escalation triggers; a hit returns the reply.Restrict
L1Confidence thresholdsRoute low-confidence replies to the lead before anything is sent.Require review
Model coreReply composed — the answer, its as-of time, triggers hit and confidence
L1 – L2Test whether a reply may be sent
L3Puts the release in the lead's hands
L4 – L5Keep the answer and its source reconstructable
L6Pulls the agent back to routing when signals degrade

How Nestack evaluates it

Evaluate the whole contact — not only the sentence that was sent.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the answer the customer reads
Depth of coverage ▼
E1Final-output evaluationDid every stated fact match the field it was read from?
E2Step-level evaluationDid the agent read the right order, the current scan and the live policy version?
E3Tool evaluationDid it read and write the correct order and the correct ticket?
E4Confidence calibrationDo low-confidence replies actually attract more corrections?
E5Slice evaluationHow does performance change across specific contact reasons?
E6Business outcomeHow many contacts reopened, escalated late or needed a correction after sending?
Floor — the outcome the retailer answers for

Failure modes

Where each failure originates in the agent

Seven failure modes, each pinned to where it starts.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
OS-03

Stale carrier scan

The last known scan is reported as the current one.

Stage gathersOrder state, carrier scans and the published policy set
02 · Reasoning2 modes
OS-04

Invented delivery date

Transit days are added to an order date on a line that never shipped.

OS-06

Warranty read as a return

A defect claim is answered with the closed return window.

Stage proposesThe answer, its as-of time and confidence
03 · Tool / write2 modes
OS-02

Address changed post-label

A labelled parcel is re-routed and the carrier record breaks.

OS-05

Cancellation noted, not called

A note lands in the desk while billing carries on unchanged.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
OS-01

Disclosure deflected

Asked whether it is a person, the agent changes the subject.

Stage returnsThe answer the customer reads and relies on
05 · Change / Version1 mode
OS-07

One channel regresses

An update drops the disclosure from voice and leaves chat intact.

Stage tracksModel, prompt, policy text and escalation config
Sev-1 · a held act performed by the agent Sev-2 · a wrong statement reaches the customer Sev-3 · source degrades, reply routes to review

Affected slices

A healthy average can rest on one contact reason

The metric is the materially-incorrect-statement rate, read as a lift against the all-contact baseline. A handful of reason codes carry most of the wrong statements Nestack reports it by slice rather than in aggregate..

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Split orders, ship date missed4.5%3.2× Review
Cancellation asked mid-contact3.5%2.5× Review
Warranty outside the window2.1%1.5× Watch
Single line, shipped, scanned0.7%0.5× Normal
Bar: materially-incorrect-statement lift vs. the all-contact baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What a cycle has to leave behind

Nothing closes on an explanation. The cycle ends when the failure exists as a regression case the next release has to pass.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Incorrect-statement rate climbs in one contact reason.

02Diagnose

The escalation lead reads the contacts and the fields behind them until one cause is left.

03Improve

The fix is written against a version, with the orders that motivated it attached.

04Verify

Release is blocked until the affected regression cases pass again.

05Learn

The case joins the permanent suite and the response playbook.

Learn → DetectThe return edge. Each pass begins with more cases to clear than the one before it.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, answering workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Contact workflow discovery and boundary definition.
02Commerce and order-system assessment.
03Policy-text, disclosure and escalation-rule mapping.
04Order and carrier ingestion with field.
05Answering logic and source-field binding.
06Confidence scoring and escalation routing.
07Escalation-lead review.
08Helpdesk and telephony integration.
09Response regression cases.
10Guardrails and escalation controls.
11Conversation-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne channel, one brand ProductionProduction contact volume AdvancedMultiple brands / regions
Introduced at Pilot
Answers from your systems of record
Escalation-lead release
Response-quality baseline
Introduced at Production
Reporting by contact reason
Escalation workflow in your desk
Approved write-back to the ticket
Helpdesk integration
Introduced at Advanced
Multi-region policy sets
Multi-stage escalation paths
High contact volume
Enterprise contact controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your contact reasons and the systems behind them Contact workflow discovery and boundary definitionWeek 1
02Representative resolved contacts Answering baseline, field binding and disclosure wordingWeek 2
03Your published shipping, return and privacy text Policy-text, disclosure and escalation-rule mappingWeek 1
04Access to relevant APIs, feeds or exports Commerce, desk and carrier assessment, then integration setupWeek 2
05Contacts you would not want repeated Response regression cases and failure-mode testingWeek 4
06Where an answer must stop and wait for a person Confidence scoring, escalation routing and approval controlsWeek 3
07A named escalation lead to review answers Escalation review workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Each phase sits on the weeks it actually occupies, and week 5 carries both evaluation and launch work.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Contact workflow discovery, policy mapping and the automation boundary W2Contact sources connected, then a baseline W3Answering workflow, confidence logic and escalation controls W4Evaluation suite, guardrails and failure-mode testing W5Helpdesk integration, pilot queues and targeted corrections W6A live contact queue answered under review, then Agent Care handover
Reading the bandEach bar covers only the weeks its work is named in. The week 5 overlap is real, not padding.
At the end of W6Validation closes on live contacts, and Agent Care picks up monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Retail AI agent

Build an order-status agent around your escalation path.

Show us your contact reasons, your order systems and who holds the exceptions. Where a contact touches a shipping promise, a warranty, a cancellation or a dispute, we will show you where the agent stops and who picks it up.

Nestack Agents · Customer service and order statusAGT-RT-01 · Agent Care available after launch