Answer internal tax questions with a drafted response, the jurisdiction and tax year it applies to and every authority cited — held for a qualified reviewer, with evaluation, audit trails and guardrails built in.
Aggregate answer quality can look acceptable while a small number of question cohorts carry most of the citation and currency failures. Nestack reports performance by slice, not only in total.
Slice performance — reported separately, not only in aggregateIllustrative example
Slice
Failure rate
Lift
Lift vs. threshold
Status
Cross-border questions
5.8%
3.4×
Review
Recent-legislation questions
3.9%
2.3×
Review
Multi-state nexus questions
2.9%
1.7×
Watch
Settled federal questions
1.0%
0.6×
Normal
Bar: lift vs. settled-question baseline · scale 0–4.0× · tick marks the 2.0× review threshold2 of 4 slices over threshold
Evidence-linked improvement
The loop does not end at Learn
Each cycle leaves the corpus re-tagged, a position re-verified and an escalation rule in force — that is what the next detection is measured against.
Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect
Citation or currency failures rise in a question cohort.
02Diagnose
Failure isolated to retrieval, authority weight, prompt or scope.
03Improve
Corpus re-tagged or an escalation rule added — approved and version-linked.
04Verify
Affected questions are re-asked against the updated corpus.
05Learn
The question becomes a standing test; the signed position joins firm precedent.
Learn → DetectThe return edge. The next cycle starts against a re-tagged corpus and one more escalation rule.
Typical build scope
Twelve workstreams across six weeks
The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, authority corpus, agent workflow, evaluation, integration, then production validation and handover.
WorkstreamWeek 1Week 2Week 3Week 4Week 5Week 6
01Workflow discovery and advice-boundary definition.
02Source-system and research-platform assessment.
03Jurisdiction, tax and question-class scoping.
04Authority corpus and effective-date tagging.
05Retrieval, citation checking and drafting.
06Confidence scoring and escalation routing.
07Reviewer sign-off workflow.
08Research and document-system integration.
09Evaluation suite and regression cases.
10Guardrails and advice-boundary controls.
11Observability and trace instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallelFinal scope and sequence confirmed in discovery
Engagement tiers
What each tier includes
Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.
Capability✓ in scope · — not at this tierPilotOne jurisdictionProductionProduction integrationAdvancedMultiple jurisdictions
Introduced at Pilot
Drafted answers with citations✓✓✓
Reviewer sign-off✓✓✓
Baseline evaluation✓✓✓
Introduced at Production
Firm precedent library—✓✓
Escalation workflow—✓✓
Approved release actions—✓✓
Observability and evaluation—✓✓
Introduced at Advanced
Multi-jurisdiction coverage——✓
Multi-stage review——✓
High query volume——✓
Enterprise controls——✓
Build priceFrom $5,000From $8,000Custom quote
Final build priceConfirmed after discovery based on jurisdictions in scope, research sources, question volume, review controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.
What we need from you
What you bring, and what we build with it
Each input maps to a piece of build scope and a week in the delivery timeline.
You bringWe build with it
01The jurisdictions, taxes and question types in scope→Jurisdiction, tax and question-class scopingWeek 1
02The advice boundary and who may sign what→Advice-boundary definition and approval controlsWeek 1
03Access to your research subscriptions and sources→Source-system and research-platform assessmentWeek 2
04Prior memos, house positions and precedent files→Authority corpus build, effective-date tagging and precedent indexingWeek 2
05Confidence and escalation thresholds→Confidence scoring, escalation routing and guardrailsWeek 3
06Questions your team has found hard to answer→Evaluation suite, regression cases and failure-mode testingWeek 4
07Named reviewers or test users→Reviewer sign-off workflow, then pilot workflow and production validationWeeks 5–6
Nothing else is requiredDeployment, documentation and Agent Care handover are ours.
Delivery timeline
Four phases across six weeks
Phases are drawn over the weeks they actually occupy. Week 5 carries both evaluation and launch work.
PhaseW1W2W3W4W5W6
DiscoveryW1
BuildW2 – W3
EvaluateW4 – W5
Pilot & LaunchW5 – W6
Week focusW1Workflow discovery, advice boundary and question-class scopingW2Authority corpus, source access and the research baselineW3Agent workflow, citation checking and reviewer controlsW4Evaluation suite, advice-boundary guardrails and failure modesW5Research-platform integration, pilot and targeted correctionsW6Sign-off on live questions, verification and Agent Care handover
Reading the bandBars cover only the weeks their work is named in. Build starts once the advice boundary is agreed in week 1.
At the end of W6Live questions are answered under sign-off, then Agent Care monitors the corpus and the escalations.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.
Next step · Accounting AI agent
Build a tax-question copilot around your research workflow.
Show us the questions your team asks, the sources you already licence and who signs what. We'll map the workflow, set the advice boundary and recommend the safest path to production.