Nestack Agent Care
Industries / Corporate Strategy / BizDev / Deep research agent

Strategy AI agent · deep research

Deep Research AI Agent

Fetch what each footnote points at before a claim enters the brief, and file the claim beside the document itself, the day it was retrieved and whether that document is primary.

4–6 weeksTypical delivery
Your stackDeployment
Retrieval-firstResearcher-held
Agent CareAfter launch

What this agent does

Opens what is cited, never rules on the claim

In
01

A claim arrives in a draft, and the citation under it is treated as an object to be fetched.

02

A footnote is offered, and it is opened, read and marked with the day it was retrieved.

Reason
03

A summary of a study is to hand, and the study itself is marked as still not held.

04

A sentence is quoted, and the words in the brief match the page or the line is marked paraphrase.

05

A document sits behind an entitlement, and the brief says where it lives and who may reach it.

Decide
06

A position taken in another country is described, and the jurisdiction it belongs to travels on.

07

A text is read in translation, and the original language and the translating hand are both named.

Out
08

A search returns nothing, and where the agent looked is written down as the finding it is.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

Retrieval, labelling and drafting belong to the agent, which sets out what it searched rather than claiming to have found everything. Each claim is accepted into the brief by a named researcher.

Example workflow

One claim, citation to acceptance

AgentHuman
1Citation evidence receivedPublic registers, licensed archives, prior briefs or documents named in a draft
2Claim context assembledThe claim, the brief it serves, the citations put forward for it and where each one now lives
3Sourced claim draftedThe claim, its citations, what each one resolved to and completeness
4Controls appliedRetrieval checks, primary-source checks, quotation checks and completeness confidence
No human action required

Stages 1 to 4 run unaided, and no claim is admitted at any of them — the agent is fetching, and a reviewer joins at the completeness gate.

5DecisionSplits at the completeness gate
Evidence sufficient

Passes to a named researcher for acceptance.

Anything thin

Adds a research lead read first.

Research review

The claim is held with its citations, what each one opened onto and the day it was fetched.

Accept · Append citation · Send to research review
Into the brief — by a named researcher
6Claim and citation records updatedOnly where write access and records policy allow it
7Outcome evaluatedRetrieval outcomes, primary-source coverage, corrections made and what the review turned up
Corrections

Each reviewer correction is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Accepting a claim into the strategy brief.
Deciding a secondary summary is good enough.
Judging a source credible on its reputation.
Signing off the brief a board will read.
Automation boundaryAgent acts unaided
Fetch every cited document that a reader could open.
Mark each source primary or secondary and say which one is in hand.
Reproduce quoted words exactly and label every paraphrase as paraphrase.
Record the retrieval date and the language read.
Nothing enters the brief and nothing is published except by a named person on approval.
Ruling on whether a claim is true.
Telling the board what the evidence means.
Setting the sourcing bar a brief must clear.
Changes to the brief or the citation record.

Example output

One claim, annotated

The market research agent we already ship serves an insights function working on commercial market figures; this one serves a strategy team, and its subject is whether a cited thing can actually be fetched and read.

Brief entry · single claimIllustrative example
Claim
Recorded as
Citation
Evidence of record
Confidence
Held for
Competitor capacity, desk research
Retrieved and read in full
Primary source
Fetched 3 August 2026
Held unaccepted
The researcher who accepts, by name
As receivedTaken from the document that was opened and dated on retrieval, and nothing beyond it is implied.
What the record holds Primary document Retrieval date Quotation checked
Why the claim waits hereWhat a retrieved document means for a decision is a judgement people make.
ActionAcceptAppend citationSend to research review
What the score decidesUnder the configured threshold the claim collects a research read before it is accepted.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every claimFrom the document behind it
03Evidence

Where the evidence is used

A brief is read long after the decision it informed, so each claim carries the day its source was retrieved, and the agent reports what it searched rather than claiming it found everything.

01Approved path

A chain ends somewhere

Walking a reference backwards stops at the document that first stated the thing, and the agent writes down where it stopped and why it could go no further.

02Human review

What was checked, and not found

One translated report was never matched to an original in the issuing archive, one working paper quoted across a sector has no publisher of record, and two licensed collections could not be opened from outside a member institution.

04Build an evidence trail

The claim, the primary source under it and the researcher who accepted it stay on file.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Public and published sourcesRegisters · archives · statistical offices
Primary documents and their dates
Licensed and subscription corporaJournal platforms · news archives
Full text behind an entitlement
Internal knowledgeDocument stores · earlier briefs
Prior claims and their sources

Agent

Deep research sourcing

Reads the brief
Fetches the sources
Holds for acceptance

Brief and record systemsResearch library · ticketing
Claim and acceptance records
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six combs between the model and the brief

Six combs drawn through in order, the last the finest. Whatever remains is set out in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeFall back to citation listing when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt and retrieval rules; a change to any of them re-opens the citation suite.Track
L4TraceabilityRecord each claim, the citations under it, what each one resolved to and every read of the brief.Record
L3Researcher releaseHold the claim for a named researcher; that hold decides release into the brief and nothing about whether the claim is true.Gate
L2Retrieval guardrailsTest each claim against the retrieval standard you configure: document opened, primary or secondary marked, quotation matched, retrieval dated.Restrict
L1Confidence thresholdsSend a claim with thin citations to a research read; a plausible reference that resolves to nothing is the worst thing this agent can hand over.Require review
Model coreEvidence assembled — the claim, its citations, what they opened onto and completeness
L1 – L2Test whether a claim may stand
L3Hands acceptance to a named person
L4 – L5Keep the claim and the source behind it
L6Returns to citation listing when signals degrade

How Nestack evaluates it

Evaluate the whole assembly — not only the claim that comes out.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the claim a brief carries
Depth of coverage ▼
E1Final-output evaluationDid the brief record what a later reader would need to open?
E2Step-level evaluationDid the agent reach the primary document, or only something written about it?
E3Tool evaluationDid the fetch succeed, and was the result stored against the right claim?
E4Confidence calibrationDo low-confidence claims actually attract more research corrections?
E5Slice evaluationHow does performance change across specific claim types?
E6Business outcomeHow many claims were sent back before a researcher would accept them?
Floor — the brief a decision rests on

Failure modes

Where each failure originates in the agent

Seven failure modes, placed where each one first becomes visible.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
ME-03

Superseded edition fetched

The version opened is no longer the current one.

Stage gathersThe briefs, the claims, the citations offered
02 · Reasoning2 modes
ME-04

Citation resolves to nothing

A reference looks right and opens on an empty page.

ME-06

Summary stands for document

A secondary account is used where the primary was needed.

Stage proposesThe claims, their citations and completeness
03 · Tool / write2 modes
ME-02

Unread citation carried on

A claim advances before its source has been opened.

ME-05

Quotation drifts from source

Words in the brief are not the words on the page.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
ME-01

Accepted, source unopened

The brief shows a claim but nothing that was read.

Stage returnsThe claim a brief carries and a reader checks
05 · Change / Version1 mode
ME-07

Silent retrieval regression

A rules change relaxes the fetch test without a note.

Stage tracksModel, prompt, retrieval rules and claim fields
Sev-1 · a claim cited to nothing at all Sev-2 · a paraphrase presented as a quote Sev-3 · source unreachable, claim held back

Affected slices

Secondary summaries absorb the corrections

A question-level retrievability figure can read clean while claims resting on secondary summaries carry most of the rework. Nestack reports the correction rate by claim type, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Claims resting on secondary summaries8.7%3.6× Review
Claims quoted from translated sources6.2%2.6× Review
Claims taken from paywalled archives3.9%1.6× Watch
Claims with a primary document in hand2.1%0.9× Normal
Bar: correction-rate lift vs. primary-document baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What an unreachable source costs

The loop shuts when the unreachable citation is a regression case. That suite is what the next brief issued is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Correction rate rises on claims resting on summaries.

02Diagnose

The footnote that resolves to a page that has never existed is taken apart until a single cause is left.

03Improve

Any change goes out numbered, with the claims that prompted it attached.

04Verify

A single failing claim case is enough to hold the release back.

05Learn

One case joins the suite, one line joins the sourcing record.

Learn → DetectThe return edge. The next brief is measured against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, claim assembly, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Retrievability-standard and automation-boundary work.
02Public, licensed and internal sources.
03Claim-to-citation and retrieval-coverage mapping.
04Citation and document ingestion.
05Claim, citation and date binding.
06Retrievability scoring and review routing.
07Researcher acceptance workflow.
08Research-library integration.
09Citation and retrieval cases.
10Guardrails and sourcing controls.
11Claim-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne claim type, one brief ProductionProduction brief workflow AdvancedMultiple briefs / languages
Introduced at Pilot
Citation retrieval to your standard
Release by a named researcher
Question-scope baseline
Introduced at Production
Reporting by research question
Brief acceptance workflow in your systems
Approved write-back
Retrieval-corpus integration
Introduced at Advanced
Multi-source briefs
Cross-language evidence packs
Large citation estates
Multi-language sourcing controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, brief volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01The briefs you produce and the questions behind them Brief inventory mapping and citation captureWeek 1
02Representative briefs, citations and source documents Citation binding, retrieval logic and the coverage baselineWeek 2
03The citation standard your briefs are held to Claim mapping, citation binding and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Public, licensed and internal corpus assessment, then integration setupWeek 2
05Briefs you would not want checked Retrieval cases and failure-mode testingWeek 4
06What no research brief may guarantee Retrievability scoring, review routing, guardrails and release controlsWeek 3
07The researcher who will accept each claim Researcher release, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Bands follow the effort a phase actually costs, and that is why one of them comes out carrying two.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Brief workflow discovery, claim mapping and the automation boundary W2Corpus integration and the retrievability baseline W3Claim assembly, retrieval logic and release controls W4Evaluation suite, retrieval cases and failure-mode testing W5Library integration, pilot claims and targeted corrections W6One research quarter run under the research lead, then Agent Care handover
Reading the bandEach bar covers only the weeks its own work is named for, and one of them has to hold a second phase.
At the end of W6Once the sourcing record validates, Agent Care assumes the agent.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Strategy AI agent

Build a deep research agent around the footnote nobody has ever opened.

Show us one claim in your last brief and the citation under it. Send that brief over and we will open every reference in it, mark which ones are primary and tell you which resolve to nothing. We set out what we searched; we do not claim to have found everything.

Nestack Agents · deep researchAGT-CS-07 · Agent Care available after launch