Rank and recommend inside the ranking policy your merchandising owner approved, serve paid placement only into slots the disclosure template owns, and hold the weights, the boosts and the surfaces for a person.
Read the query or session, the catalogue, live inventory and the ranking config from supported search.
02
Standardise queries, synonyms and product attributes, and carry each ranking input with the source it came from.
Reason
03
Score and order candidates using the approved policy and the feature set signed off alongside it.
04
Separate the slots — which are organic, which are monetised, which the disclosure template owns.
05
Attach to each surfaced result the inputs that lifted it, commercial signals included.
Decide
06
Mark any result whose position rests on a paid, supplier-funded or own-brand signal.
07
Send ranking-policy and surface changes to the category merchandising manager rather than applying them.
Out
08
Retain the query, the ranking inputs, the disclosure state and the set that was served.
09
Execute write actions only inside the approval boundaries agreed during implementation.
→Product statement
The agent re-ranks inside an approved policy; the category merchandising manager owns the weights, the boost rules and which surfaces carry paid placement at all.
Example workflow
One session, query to served set
AgentHuman
1Query or session receivedOn-site search box, category page, personalised carousel or app surface
2Signals assembledCatalogue records, live inventory, eligibility state and the signed ranking config
3Ranking proposedOrdered results, slot assignment, commercial-signal contribution and confidence
4Gates appliedConfig-version checks, feature-set checks, slot and disclosure checks and confidence threshold
No human action required
Stages 1 to 4 run unaided and no ranking policy moves at any of them — the merchandising lane opens at the confidence gate.
5DecisionSplits at the disclosure gate
Clearly relevant
Serves inside the approved config.
Ambiguous intent
Takes an ad-operations read first.
Merchandising approval
The ranking is held with its inputs, its marked slots and the confidence.
Approve · Adjust · Send to ad ops
Approved — released to serve▼
6Search and personalisation systems updatedOnly where write access and approval policy allow it
Integration availability depends on the client's existing systems and API access.
Agent controls
Six layers between the model and the shopper
Every layer wraps the next. What none of them catches is in the map below.
L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeRoll the agent back to the last approved ranking config when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, ranking-config and feature-set changes.Track
L4TraceabilityRecord the query, the inputs, the slot assignment, the disclosure state and the approver.Record
L3Merchandising approvalHold policy and surface changes for the named merchandising manager; it governs release, not whether an approved ranking is right.Gate
L2Policy guardrailsTest rankings against the approved feature set, the slot rules and the disclosure template; a failure returns the ranking.Restrict
L1Confidence thresholdsRoute low-confidence rankings to an ad-operations read before they are served.Require review
Model coreRanking produced — ordered results, slot assignment, commercial signals and confidence
L1 – L2Test whether a ranking may stand
L3Puts the release in a merchant's hands
L4 – L5Hold the inputs the ranking rested on
L6Falls back to the default ordering when signals slip
How Nestack evaluates it
Evaluate the whole ranking path — not only the top result.
Coverage runs the whole depth of the workflow, and every layer is cut by slice.
Surface — the set the shopper sees
Depth of coverage ▼
E1Final-output evaluationWas this the order the approved policy produces, with each label rendered?
E2Step-level evaluationDid the agent use the signed config, the approved feature set and live inventory?
E3Tool evaluationDid it read and write the correct index and the correct surface?
E4Confidence calibrationDo low-confidence rankings actually attract more merchandising adjustments?
E5Slice evaluationHow does performance change across specific query classes?
E6Business outcomeHow many impressions carried a commercial placement with no rendered label?
Floor — what the shopper sees ranked
Failure modes
Where each failure originates in the agent
Seven ways the ranking goes wrong, by stage.
Agent lifecycleDirection of processing →
01 · Retrieval1 mode
RK-03
Stale monetisation manifest
A supplier-funded item is scored as organic and served unlabelled.
Stage gathersCatalogue, inventory, with the source each came from
02 · Reasoning2 modes
RK-04
Proxy feature promoted
A location-derived affluence signal enters the score as a proxy.
RK-06
Review set trimmed by rating
Low-star reviews are down-weighted while the summary reads as complete.
Stage proposesOrdered results, slot assignment and confidence
03 · Tool / write2 modes
RK-02
Config written past sign-off
A new ranking config reaches the feature store before merchandising sees it.
RK-05
Cohort enrolled without consent
A mixed-audience surface adds sessions with no parental consent on file.
Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
RK-01
Disclosure not visible
The sponsored label renders in low contrast and below the fold.
Stage returnsThe set the shopper sees and the labels on it
05 · Change / Version1 mode
RK-07
Parameters move, text does not
An upgrade shifts the main ranking parameters while the published explanation still describes the last set.
Stage tracksModel, prompt, ranking config and feature set
Sev-1 · a held act performed by the agentSev-2 · an unlabelled paid placement is servedSev-3 · signals degrade, ranking routes to review
Start at the all-session baseline: across ordinary traffic the undisclosed-commercial-influence rate sits low. Read the same rate by cohort and a few surfaces carry several times.
Slice performance — reported separately, not only in aggregateIllustrative example
Slice
Failure rate
Lift
Lift vs. threshold
Status
Mixed-audience sessions, sponsored slots
4.1%
3.7×
Review
EU and UK regulated categories
3.1%
2.8×
Review
Logged-out long-tail queries
2.0%
1.8×
Watch
Logged-in head queries, organic only
1.0%
0.9×
Normal
Bar: undisclosed-commercial-influence lift vs. the all-session baseline · scale 0–4.0× · tick marks the 2.0× review threshold2 of 4 slices over threshold
Evidence-linked improvement
A cycle is measured by what it adds
The loop shuts when the miss is a case in the suite, not when someone has explained it.
Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect
Undisclosed placements cluster on one surface.
02Diagnose
The category merchandising manager reads the impressions and the configs behind them, and does not stop at the first plausible cause.
03Improve
Whatever changes ships against a version, with the queries that prompted it attached.
04Verify
Nothing ships until the affected cases pass a second time.
05Learn
The suite grows by one case; so does the ranking policy.
Learn → DetectThe return edge. What is detected next is measured against a suite one case longer.
Typical build scope
Twelve workstreams across six weeks
The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, ranking workflow, evaluation, integration, then production validation and handover.
WorkstreamWeek 1Week 2Week 3Week 4Week 5Week 6
01Ranking workflow discovery and boundary definition.
02Search and catalogue assessment.
03Ranking-policy, slot and disclosure-rule mapping.
04Catalogue and session ingestion.
05Ranking logic and input binding.
06Confidence scoring and routing.
07Merchandising review workflow.
08Commerce-platform integration.
09Relevance regression cases.
10Guardrails and surfacing controls.
11Session-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallelFinal scope and sequence confirmed in discovery
Engagement tiers
What each tier includes
Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.
Capability✓ in scope · — not at this tierPilotOne surface, one marketProductionProduction search trafficAdvancedMultiple markets / brands
Introduced at Pilot
Ranking to your approved policy✓✓✓
Merchandising approval✓✓✓
Relevance baseline✓✓✓
Introduced at Production
Reporting by query class—✓✓
Approval workflow in your systems—✓✓
Approved config write-back—✓✓
Search-analytics integration—✓✓
Introduced at Advanced
Multi-region ranking rules——✓
Multi-stage merchandising approvals——✓
High query volume——✓
Enterprise ranking controls——✓
Build priceFrom $5,000From $8,000Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, transaction volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.
What we need from you
What you bring, and what we build with it
Each input maps to a piece of build scope and a week in the delivery timeline.
You bringWe build with it
01Your ranking policy and its boost rules→Ranking-policy, slot and disclosure mappingWeek 1
02Representative queries and sessions→Ranking baseline, input binding and slot rulesWeek 2
03Your surface map, organic and monetised→Surface and disclosure mapping, and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports→Search, personalisation and retail-media assessment, then integration setupWeek 2
05Queries that surfaced the wrong thing→Relevance cases and the evaluation suiteWeek 4
06Which results may never be reordered automatically→Confidence bands, surfacing rules and review controlsWeek 3
07Named merchandisers to review the ranking→Merchandising approval workflow, then pilot and production validationWeeks 5–6
Nothing else is requiredDeployment, documentation and Agent Care handover are ours.
Delivery timeline
Four phases across six weeks
The bands follow real work rather than a plan, so evaluation and pilot genuinely share the fifth week.
PhaseW1W2W3W4W5W6
DiscoveryW1
BuildW2 – W3
EvaluateW4 – W5
Pilot & LaunchW5 – W6
Week focusW1Ranking workflow discovery, surface mapping and the automation boundaryW2Catalogue and session data connectedW3Ranking logic, confidence scoring and surfacing controlsW4Relevance testing and disclosure guardrailsW5Retail-media integration, pilot surfaces and targeted correctionsW6A live ranking release under merchandising, then Agent Care handover
Reading the bandNothing is padded to fill a week. Evaluation and pilot both land in week 5.
At the end of W6Once the release validates, Agent Care owns the running agent.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.
Next step · Retail AI agent
Build a recommendation agent around the policy your merchandiser already signs.
Show us your surfaces, your ranking policy and where paid placement sits today. Then the question the build turns on: which slots may an agent reorder unattended, and which ones does the merchandising owner keep?