Forecast at the level you actually order at, publish every forecast beside the simple baseline it had to beat, and hand the named planner the error they will own.
Ingest sales history, prices, promotion calendars and event dates from planning, ERP or POS sources.
02
Normalise item, store and calendar fields, and carry each series forward with its own history.
Reason
03
Forecast at the item-store level the ordering decision is made at, not only in aggregate.
04
Run the naive and seasonal-naive baselines alongside, on the same holdout, over the same period.
05
Report the forecast beside those baselines, so the comparison is one a planner can re-run.
Decide
06
Flag new products and promotional periods as regimes the forecast does not cover.
07
Route every forecast to the named demand planner, who decides what enters the plan.
Out
08
Retain the history, the baseline run, the planner's overrides and the forecast accepted.
09
Execute write actions only inside the approval boundaries agreed during implementation.
→Product statement
The agent proposes a forecast against a benchmark; the demand planner accepts it into the plan, and ordering, promotion and production stay separate decisions.
Example workflow
One item, history to accepted forecast
AgentHuman
1History receivedPOS, ERP or planning-system history, prices and the promotion calendar
2Context assembledPrice, promotion, event and calendar features, each with the series it belongs to
Integration availability depends on the client's existing systems and API access.
Agent controls
Six layers between the model and the plan
Each control encloses the last. Whatever slips past all of them is named below.
L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeCut back to the naive baseline when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, feature-set and horizon configuration changes.Track
L4Forecast trailRecord the history, the baseline run, the overrides and the acceptance.Record
L3Planner acceptanceHold forecasts for a named planner; acceptance governs entry to the plan, not whether the number is right.Gate
L2Benchmark guardrailTest each forecast against the naive and seasonal-naive baselines on the same holdout; losing to one returns the forecast.Restrict
L1Confidence thresholdsRoute low-confidence forecasts back to the baseline before the planner sees them.Require review
Model coreForecast produced — horizon, baseline run, regime flags and confidence
L1 – L2Test whether a forecast may stand
L3Puts the acceptance in a planner's hands
L4 – L5Keep the forecast and the history behind it
L6Cuts back to the naive baseline when signals degrade
How Nestack evaluates it
Evaluate the forecasting workflow — not only the number at the end.
Coverage runs the whole depth of the workflow, and every layer is cut by slice.
Surface — the forecast the plan sees
Depth of coverage ▼
E1Final-output evaluationDid the forecast beat its stated benchmark on the same holdout?
E2Step-level evaluationDid the agent use the right history, calendar and feature set?
E3Tool evaluationDid it read and write the correct item and the correct store?
E4Confidence calibrationDo low-confidence forecasts actually attract more planner overrides?
E5Slice evaluationHow does performance change across specific item groups?
E6Business outcomeHow many forecasts needed an override, or missed what the plan needed?
Floor — the plan the business answers for
Failure modes
Where each failure originates in the agent
Seven ways a forecast goes wrong, placed by stage.
Agent lifecycleDirection of processing →
01 · Retrieval1 mode
BL-03
Superseded history read
A restated sales series is read as though it were current.
Stage gathersSales history, prices, promotions and event dates
02 · Reasoning2 modes
BL-04
Promotion read as normal
A promoted week is forecast as though it were ordinary trading.
BL-06
New product proxied
An item with no history is forecast from a proxy nobody chose.
Stage proposesForecast, baseline run, regime flags and confidence
03 · Tool / write2 modes
BL-02
Ordered before acceptance
A forecast reaches an order before a planner accepted it.
BL-05
Duplicate series
One item-store series is forecast twice from two sources.
Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
BL-01
Beaten by the baseline
The forecast is worse than the simple method beside it.
Stage returnsThe forecast the planner accepts into the plan
05 · Change / Version1 mode
BL-07
Silent error regression
A model or feature change widens the error unannounced.
Stage tracksModel, prompt, feature sets and horizon config
Sev-1 · an order ran off a forecastSev-2 · the baseline beat the forecastSev-3 · history degrades, forecast to review
Aggregate override rates can look acceptable while a few item cohorts absorb most of the rework. Nestack reports the planner-override rate by slice, not only in total.
Slice performance — reported separately, not only in aggregateIllustrative example
Slice
Failure rate
Lift
Lift vs. threshold
Status
New products with no history
5.6%
3.2×
Review
Promotional periods
4.4%
2.5×
Review
Slow-moving long-tail items
2.4%
1.4×
Watch
Established weekly sellers
1.6%
0.9×
Normal
Bar: planner-override-rate lift vs. established-seller baseline · scale 0–4.0× · tick marks the 2.0× review threshold2 of 4 slices over threshold
Evidence-linked improvement
The loop ends in a test, not a review
The cycle does not end in a post-mortem. It ends in a case the next release has to clear, and that suite is what the next forecast issued is measured against.
Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect
Override rate rises in an item slice.
02Diagnose
The forecasts come out first — the series, the benchmark run beside each one and the overrides on top — and are read until the cause narrows to one.
03Improve
Each change carries a version and the forecasts that forced it.
04Verify
Release waits on the affected cases going green again.
05Learn
The case is kept for good, and the planning rules are rewritten.
Learn → DetectThe return edge. The next forecast meets a suite one case longer.
Typical build scope
Twelve workstreams across six weeks
The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, data, forecasting workflow, evaluation, integration, then production validation and handover.
WorkstreamWeek 1Week 2Week 3Week 4Week 5Week 6
01Demand-planning workflow discovery and boundary definition.
02Planning, ERP and POS source assessment.
03Baseline, holdout and error-metric agreement with planning.
04History ingestion and normalisation.
05Forecast logic and baseline binding.
06Confidence scoring and regime routing.
07Planner acceptance workflow.
08Planning-system integration.
09Benchmark and promotion cases.
10Guardrails and override controls.
11Forecast-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallelFinal scope and sequence confirmed in discovery
Engagement tiers
What each tier includes
Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.
Capability✓ in scope · — not at this tierPilotOne category, one bannerProductionProduction planning systemsAdvancedMultiple banners / markets
Introduced at Pilot
Forecasting to your history and benchmarks✓✓✓
Planner acceptance✓✓✓
Benchmark-comparison baseline✓✓✓
Introduced at Production
Reporting by category—✓✓
Acceptance workflow in your systems—✓✓
Approved write-back—✓✓
Planning-system integration—✓✓
Introduced at Advanced
Multi-source demand signals——✓
Multi-stage planning approvals——✓
High item-store counts——✓
Multi-banner forecast controls——✓
Build priceFrom $5,000From $8,000Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, item-store counts, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.
What we need from you
What you bring, and what we build with it
Each input maps to a piece of build scope and a week in the delivery timeline.
You bringWe build with it
01Your item master and sales history→History ingestion and item mappingWeek 1
02Representative seasons and promotions→Forecast baseline, benchmark runs and holdout designWeek 2
03The baseline you run today and your error metric→Baseline, holdout and error-metric agreement with planningWeek 1
04Access to relevant APIs, feeds or exports→Planning, ERP and POS assessment, then integration setupWeek 2
05Forecasts you would not want ordered against→Benchmark cases and failure-mode testingWeek 4
06Where a forecast must stop and wait for a planner→Confidence scoring, regime routing, guardrails and override controlsWeek 3
07Named demand planners to accept forecasts→Planner acceptance workflow, then pilot and production validationWeeks 5–6
Nothing else is requiredDeployment, documentation and Agent Care handover are ours.
Delivery timeline
Four phases across six weeks
The bands sit on the weeks the work occupies, which is why the fifth carries two kinds at once.
PhaseW1W2W3W4W5W6
DiscoveryW1
BuildW2 – W3
EvaluateW4 – W5
Pilot & LaunchW5 – W6
Week focusW1Demand-planning discovery, benchmark agreement and the boundaryW2Source integration and the forecast baselineW3Forecasting workflow, confidence logic and acceptance controlsW4Benchmark cases, held-act guardrails and failure-mode testingW5Planning-system integration, pilot categories and targeted correctionsW6One planning cycle forecast under the demand team, then handover
Reading the bandA bar covers only the weeks its work is named in; the fifth carries two kinds at once.
At the end of W6The cycle closes the validation and Agent Care takes over monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.
Next step · Food & Beverage AI agent
Build a forecasting copilot that publishes its benchmark.
Show us your sales history, your promotion calendar and the method you forecast with today. What does that method beat, at the level you actually order at, and who owns the error when it misses? We build the boundary around that answer.