Nestack Agent Care

Retail AI agent · Payment fraud

Payment-Fraud & Chargeback-Recovery AI Agent

Score the transaction, hold the risky ones for a short review rather than declining outright, and leave every decline above the threshold to the fraud analyst who decides it.

4–6 weeksTypical delivery
Your stackDeployment
Hold, not blockAnalyst decides
Agent CareAfter launch

What this agent does

Scores the order, not the person behind it

In
01

Ingest orders, authentication results and payment signals from supported commerce, gateway and risk sources.

02

Normalise the fields, and carry every signal forward with the source it came from.

Reason
03

Score the transaction and rank it against the review queue.

04

Apply the rules, thresholds and review policy configured for the merchant.

05

Keep the score about the transaction, and exclude protected characteristics and their close proxies.

Decide
06

Word any customer message about the order alone, not as a claim about the customer.

07

Route every decline above the threshold to the fraud analyst who decides it.

Out
08

Retain the score, its top features, the model version, the rules fired and the reviewer.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

The agent scores; the fraud analyst decides — and a wrongly declined order is a customer who simply leaves, while a wrongly filed representment is a network-rule matter with its own limits.

Example workflow

One order, score to decision

AgentHuman
1Order receivedA checkout, a stored-credential reorder or a card-not-present transaction
2Features assembledAuthentication result, payment signals, order history and tenure, each with its source
3Transaction scoredThe score, its top contributing features, the rules fired and confidence
4Risk checks appliedExcluded-feature checks, threshold policy, message-wording rules and confidence threshold
No human action required

Stages one to four run unaided — the agent scores and ranks, and the analyst's lane opens at the confidence gate for every decline that has to be reviewed.

5DecisionSplits at the decision threshold
Low risk

Goes to the fraud analyst to decide.

Elevated risk

Adds a senior risk read first.

Fraud-analyst review

The order is held with its score, its top features and the confidence.

Release · Uphold · Escalate to risk
Decided — released or declined
6Order systems updatedOnly where write access and approval policy allow it
7Decision evaluatedFalse-decline rate, network programme ratios, representment outcomes and customer appeals
Overturns

Every analyst release is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Declining any order above the review threshold.
Adding any customer to a permanent block list.
Setting a score threshold or a fraud rule.
Approving a representment packet before submission.
Automation boundaryAgent acts unaided
Score transactions and rank them for the analyst review queue.
Place a short, time-boxed hold pending that review.
Assemble and verify representment evidence.
Monitor network programme ratios and escalate as they approach.
Any write happens inside the boundaries agreed at implementation, never ahead of approval.
Wording of a customer-facing decline or dispute message.
Deploying a model version and signing its disparity test.
Onboarding a data source and classifying it for consumer-report status.
Remediation and notice posture after a wrongful decline.

Example output

One decision, annotated

Everything the agent proposes is attached to the signals it was drawn from.

Fraud output · single orderIllustrative example
Customer
Agent action
Order value
Top contributing feature
Confidence
Customer wording
Repeat domestic customer
Held for analyst review rather than declined outright
$412.00
New shipping address
88%
Transaction-neutral only
As receivedTaken from the order, the authentication record and this customer's own history with the merchant.
Signals used Authentication result Order and tenure history Device and address match
Why this holdThe score is about the transaction, not the customer and a person still decides.
ActionReleaseUpholdEscalate to risk
What the score decidesBeyond the configured threshold the hold goes to the fraud analyst rather than standing.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every transactionFrom the checkout and the gateway
03Scoring

Score the transaction

Draw on authentication results, payment signals and the customer's own history with the merchant.

01Approved path

Decline less, review better

Routine low-risk orders clear without an analyst touching them.

02Human review

Send review to the costly calls

High-value, low-tenure and cohort-flagged orders are marked, so the analyst's read starts where a wrong call costs most.

04Build an evidence trail

The score, its inputs and the analyst who reviewed it stay with the decision.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Gateway-native fraud scoringStripe Radar · Adyen RevenueProtect
Sift · ClearSale
Fraud decisioning platformsSignifyd · Riskified
Forter · Kount
Dispute and representmentChargebacks911 · Justt
Midigator · Chargeflow

Agent

Payment fraud & chargeback recovery

Reads the order
Scores the transaction
Holds for review

Pre-dispute and workflowVerifi · Ethoca
Acquirer and gateway APIs
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the decline

The stack nests inward; the map below names what each layer misses.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modePull scoring back to ranking-only when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, threshold and rule-set changes.Track
L4TraceabilityRecord the score, its top features, the model version, the rules fired and the reviewer.Record
L3Fraud-analyst reviewHold declines for the named analyst; it governs the decision, not whether the score was right.Gate
L2Policy guardrailsTest decisions against excluded features, threshold policy and message wording; a failure returns the case.Restrict
L1Confidence thresholdsRoute low-confidence scores to a senior risk read before the analyst decides.Require review
Model coreOrder scored — the score, its top features, the rules fired and confidence
L1 – L2Test whether a decision may stand
L3Puts the decline in an analyst's hands
L4 – L5Keep the score and its inputs reconstructable
L6Widens review rather than declining more

How Nestack evaluates it

Evaluate the decision workflow — not only the fraud that was caught.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the outcome the customer meets
Depth of coverage ▼
E1Final-output evaluationWas the decline supported by transaction evidence rather than by who the customer is?
E2Step-level evaluationDid the agent use the right signals, thresholds and review policy?
E3Tool evaluationDid it write the correct order and the correct hold state?
E4Confidence calibrationDo low-confidence scores actually attract more analyst releases?
E5Slice evaluationHow does performance change across specific customer segments?
E6Business outcomeHow many legitimate orders were declined, held to abandonment or wrongly represented?
Floor — the transaction that stands or falls

Failure modes

Where each failure originates in the agent

Seven modes across the decision lifecycle.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
FD-03

Stale reputation record

A recycled address carries a prior ring's history onto a new customer.

Stage gathersOrders, authentication results, with the source each came from
02 · Reasoning2 modes
FD-04

Gift order read as fraud

A billing and shipping mismatch on a gift is scored as risk.

FD-06

Eligibility miscounted

Transactions outside the qualifying window are counted as evidence.

Stage proposesThe score, its top features and confidence
03 · Tool / write2 modes
FD-02

Permanent block written

A time-boxed hold is written as a standing block across systems.

FD-05

Wrong evidence attached

A representment carries a delivery record from another order.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
FD-01

Accusation in the message

A decline notice tells the customer their order was fraudulent.

Stage returnsThe message and the decision the customer meets
05 · Change / Version1 mode
FD-07

Silent calibration shift

A model refresh moves the score scale under a fixed threshold.

Stage tracksModel, prompt, thresholds and rule sets
Sev-1 · a decision made outside the boundary Sev-2 · a legitimate order is declined Sev-3 · a signal degrades

Affected slices

Overall approval rates can hide one segment

The false-decline rate is the share of legitimate orders declined, or held until the customer abandons them. Nestack reports it by slice as a lift on the all-transaction baseline.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
First-time, forwarder, high value6.5%3.8× Review
Consumer-report identity checks4.8%2.8× Review
Established customers, real victims2.4%1.4× Watch
Repeat domestic, stored credential1.2%0.7× Normal
Bar: false-decline lift vs. the all-transaction baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

A cycle closes on a case, not a review

A cycle finishes when the failure has a case, an owner and a release that must pass it.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

False declines rise in one customer segment.

02Diagnose

While the gap in that segment stays open, the analyst replays the declines and their features until one cause is left.

03Improve

The correction is stamped and linked to the decisions that produced it.

04Verify

Re-running the affected cases is the release condition.

05Learn

A permanent case, plus a change to the decision policy.

Learn → DetectThe return edge. Each detection meets more standing cases than the one before.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, decision workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Decision workflow discovery and boundary definition.
02Gateway and processor-source assessment.
03Rule, threshold and review-policy mapping and rule mapping.
04Order and payment-record ingestion.
05Scoring logic and signal-provenance.
06Confidence scoring and review routing.
07Fraud-analyst review.
08Gateway, risk and dispute integration.
09Decision regression cases.
10Guardrails and decline controls.
11Score-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne channel, one payment set ProductionProduction decision volume AdvancedMultiple channels / regions
Introduced at Pilot
Scoring to your rules and thresholds
Fraud-analyst review
Decision-quality baseline
Introduced at Production
Reporting by segment
Review workflow in your systems
Approved write-back to order systems
Risk-platform integration
Introduced at Advanced
Multi-region decision rule sets
Multi-stage risk approvals
High transaction volume
Enterprise risk controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your decline rules, thresholds and review policy Rule, threshold and review-policy mappingWeek 1
02Representative past decisions Scoring baseline, feature review and provenance bindingWeek 2
03Your vendors and their consumer-report status Threshold and review-policy mapping, and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Gateway, risk and dispute assessment, then integration setupWeek 2
05Declines you had to apologise for Decision cases and failure-mode testingWeek 4
06Which declines must be reviewed before they stand Risk bands, review routing and decline controlsWeek 3
07A named analyst to review declines Analyst review workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Bands follow the work rather than the calendar, and evaluation shares week 5 with the pilot.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Decision workflow discovery, threshold mapping and the automation boundary W2Signal and network data connected W3Scoring workflow, confidence logic and review controls W4Decision cases and decline guardrails W5Dispute integration, pilot decisions and targeted corrections W6A live decision window reviewed by the analyst, then handover
Reading the bandNo padding — week 5 genuinely carries both evaluation and pilot.
At the end of W6Checks close on live decisions and Agent Care takes the monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Retail AI agent

Build a fraud agent that is precise about which law actually applies.

A merchant declining a card is not a creditor taking adverse action; the real hooks are state unfair-practices and public-accommodation law, and fair-credit-reporting duties where a consumer report is used. Show us your decline rules and who signs a representment.

Nestack Agents · Payment fraud and chargeback recoveryAGT-RT-13 · Agent Care available after launch