Nestack Agent Care

Insurance AI agent · AI governance

AI-Governance & Market-Conduct Evidence AI Agent

Keep the AI-systems inventory current — built and bought — hold the governance record each state expects, track when testing is due, and assemble the evidence pack an examiner asks for.

4–6 weeksTypical delivery
Your stackDeployment
Inventory-ledNamed officer
Agent CareAfter launch

What this agent does

Keeps the inventory, never the attestation

In
01

A model is registered, and the agent logs its type, owner and risk classification into the inventory.

02

A vendor system arrives under contract, and the agent adds it to the same inventory as anything built in-house.

Reason
03

The state-required governance record is assembled against that entry: risk assessment, testing and review date.

04

Vendor due-diligence files and contract terms are checked for the audit rights the bulletin requires.

05

Testing and review due dates are tracked against every state and line a system touches.

Decide
06

Systems missing documentation, a due-diligence gap or an overdue test are flagged for the officer.

07

The flagged items and the assembled record are routed to the named compliance officer to review.

Out
08

The inventory, the governance record, the vendor file and the officer's review stay attached to the system.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

The agent assembles the inventory and the evidence; the named compliance officer decides what the record shows and is the one who attests.

Example workflow

One system, inventory to evidence

AgentHuman
1System identifiedModel registration, vendor contract, deployment request or existing inventory record
2Inventory recordedModel type, owner, risk classification, states and lines in use, and vendor terms
3Governance record assembledRisk assessment, testing history, review date, audit-rights status and confidence
4Controls appliedDue-diligence checks, testing-cadence checks, state-scope checks and confidence threshold
No human action required

Stages 1 to 4 run unaided, and nothing is attested at any of them — the agent is assembling, and the officer's lane opens at the confidence gate.

5DecisionBranches at the confidence threshold
High confidence

Goes to the officer's queue as a completed record.

Low confidence

Routes for a documentation gap to close before the officer reviews it.

Officer review

The record is held with its evidence, its flagged gaps and the confidence.

Attest · Request more evidence · Escalate
Attested — released to the record
6Inventory and governance systems updatedOnly where write access and approval policy allow it
7Outcome evaluatedDocumentation completeness, flag outcomes, testing-cadence compliance and post-attestation corrections
Evidence requests

Every evidence request is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Signing or submitting a regulatory attestation.
Certifying compliance with any bulletin or regulation.
Concluding a model is fair or unbiased.
Approving a system for production use.
Automation boundaryAgent acts unaided
Maintain the inventory of every AI system in use, built and bought.
Assemble the governance record each state's rule requires.
Track when testing and review are due, by system and by state.
Flag documentation and audit-rights gaps, and hold.
Any write happens inside the boundaries agreed at implementation, never ahead of the officer.
Answering a regulator or examiner directly.
Treating a vendor's proprietary-process claim as documentation.
Classifying a system's risk tier without review.
Changes to inventory rules or state-scope configuration.

Example output

One entry in the inventory, annotated

Everything the agent assembles is attached to the record it came from.

Governance-record output · single systemIllustrative example
System
Flagged status
Testing status
Evidence type
Confidence
Attribution
Vendor-supplied fraud-scoring tool
Logged as internally built; the vendor contract on file says otherwise
212 days since last test
Vendor due-diligence file
81%
Compliance officer of record
As receivedTaken from the inventory record and the vendor file on hand — nothing on this side is decided by the agent.
Evidence used Vendor contract terms Prior testing log State scope record
Why this is flaggedThe build classification conflicts with the contract on file — the line the officer must resolve.
ActionAttestRequest more evidenceEscalate
What the score decidesBelow the configured threshold before it reaches the approver.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every systemFrom the inventory record
03Assembly

Assemble against the record

Draw on the inventory entry and the state governance rules configured for that system.

01Approved path

Bought is still yours to explain

A system bought from a vendor sits on the insurer's inventory the same as one built in-house — the record has to say so.

02Human review

Point the officer at

Missing due-diligence files, an expired test date or an unclear risk tier are flagged, so review starts where the gap sits.

04Build an evidence trail

The system, the testing run against it and the officer who owns it stay in the inventory.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Model inventory & MLOpsModel registries · ML platforms
MLflow · Databricks · SageMaker
Vendor & contract recordsDue-diligence and contract files
DocuSign CLM · Ironclad · Coupa
GRC & compliance systemsPolicy and control platforms
ServiceNow GRC · OneTrust · Archer

Agent

AI-governance & market-conduct evidence

Reads the inventory
Assembles the record
Holds for the officer

Regulatory correspondenceData calls · exam requests
State DOI portals · exam tools
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the record

Every control wraps the one within it. What the set does not catch is named in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modePull the agent back to inventory reporting only when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt, rule and state-configuration changes.Track
L4TraceabilityRecord the system, its evidence, the flags and the officer's review.Record
L3Officer reviewHold the record for the named officer; it governs release, not whether the record is complete.Gate
L2Policy guardrailsTest the record against state-scope and documentation rules; a failure returns it.Restrict
L1Confidence thresholdsRoute low-confidence assemblies for more evidence before the officer sees them.Require review
Model coreRecord assembled — inventory entry, governance record, flagged gaps and confidence
L1 – L2Test whether the record may proceed
L3Puts the release in the officer's hands
L4 – L5Keep the system and the testing behind it
L6Reverts to inventory reporting when signals degrade

How Nestack evaluates it

Evaluate the governance workflow — not only the final record.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the record the examiner sees
Depth of coverage ▼
E1Final-output evaluationDid every entry in the record match the evidence on file?
E2Step-level evaluationDid the agent apply the right state rules and testing cadence for the system?
E3Tool evaluationDid it read and write the correct system and the correct field?
E4Confidence calibrationDo low-confidence assemblies actually need more evidence?
E5Slice evaluationHow does performance change across specific system types?
E6Business outcomeHow many records needed more evidence or a later correction?
Floor — the outcome the insurer answers for

Failure modes

Where each failure originates in the agent

Seven failure modes, placed at the stage each one originates.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
GO-03

Stale system record

A system's testing or risk-tier status is read from a record already superseded.

Stage gathersInventory entries, vendor files and testing logs
02 · Reasoning2 modes
GO-04

Vendor system excluded

A bought system is left off the inventory as not the insurer's own.

GO-06

EU clocks conflated

The transparency date and the high-risk date are merged into one deadline.

Stage proposesInventory entry, evidence, flags and confidence
03 · Tool / write2 modes
GO-02

Attestation drafted unreviewed

A governance summary reads as a fairness attestation before the officer opens it.

GO-05

Retirement deletes the trail

A retired system's testing history is deleted instead of archived.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
GO-01

Audit-rights gap surfaces late

A vendor contract lacks audit rights, unnoticed until a data call asks for the report.

Stage returnsThe record the officer attests to and answers for
05 · Change / Version1 mode
GO-07

Silent scope drift

A rule or model change widens what counts as adequate due diligence.

Stage tracksModel, prompt, state rules and configuration changes
Sev-1 · a write acts outside the boundary Sev-2 · unreviewed claim reaches the record Sev-3 · evidence degrades, routes to review

Affected slices

A clean score can hide one bad system

A group-wide documentation score can look complete while a handful of AI-system cohorts carry most of the gap. Nestack reports the documentation-gap rate by slice, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Third-party systems bought, not built7.9%4.0× Review
Models changed since the last review5.6%2.9× Review
Systems used in more than one state3.3%1.7× Watch
Internally built, documented at release1.9%0.8× Normal
Bar: documentation-gap rate lift vs. the internally-built baseline · scale 0–4.0× · tick at 2.0× 2 of 4 slices over threshold

Evidence-linked improvement

Every cycle closes with one more documented case

The loop shuts when the undocumented model is a regression case Each cycle leaves the next one a harder test to pass..

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Documentation-gap rate rises in a system slice.

02Diagnose

The model nobody wrote down is traced back through the inventory until the cause narrows to one gap.

03Improve

Whatever changes ships against a version, with the systems that prompted it attached.

04Verify

Nothing ships until the affected documentation cases pass a second time.

05Learn

The suite grows by one case; so does the evidence pack.

Learn → DetectThe return edge. Detection next time runs against a suite one system longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, governance workflow, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Governance workflow discovery and scope and boundary definition.
02Inventory and vendor-source assessment.
03State-rule and documentation mapping and rule mapping.
04Inventory ingestion and normalisation.
05Record-assembly logic and source binding.
06Confidence scoring and gap routing.
07Officer review workflow.
08Vendor and GRC system integration.
09Inventory and documentation cases.
10Guardrails and attestation controls.
11System-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne business unit, one state ProductionProduction governance systems AdvancedMultiple states / lines
Introduced at Pilot
Inventory and record assembly to your rules
Officer review
Documentation-completeness baseline
Introduced at Production
Reporting by state
Officer workflow in your systems
Approved write-back
Model-inventory integration
Introduced at Advanced
Multi-state governance rules
Multi-stage officer approvals
High system volume
Multi-state governance controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, system volume, review controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your AI-system inventory and vendor-contract list Inventory ingestion and vendor-contract mappingWeek 1
02Representative past systems, including retired ones Governance-record assembly baseline and testing-cadence logicWeek 2
03Your state-scope list and documentation policy State-rule and documentation-policy mappingWeek 1
04Access to relevant APIs, feeds or exports Inventory and GRC-system assessment, then integration setupWeek 2
05Systems you would not want examined Data-call cases and failure-mode testingWeek 4
06What no system may go undocumented on Confidence scoring, gap routing, guardrails and attestation controlsWeek 3
07Named compliance officers to review records Officer review workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Bands follow the real work rather than the plan, which is why evaluation and pilot share week 5.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Governance workflow discovery, state-rule mapping and the automation boundary W2Inventory and vendor-source integration and the assembly baseline W3Assembly workflow, confidence logic and officer-review controls W4Evaluation suite, documentation checks and failure-mode testing W5Vendor-system integration, pilot systems and targeted corrections W6One reporting cycle assembled under the compliance officer, then handover
Reading the bandA bar covers the weeks its work is named in, and nothing else. The week 5 overlap is real, not padding.
At the end of W6Once the cycle validates, Agent Care owns the running agent.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Insurance AI agent

Build an AI-governance agent — one your own inventory will have to list, too.

Bought is still yours to explain, and the record has to say so. Show us your AI-system inventory, your vendor contracts and the officer who signs off — we'll map what's tracked, what's missing and where the documentation gap sits.

Nestack Agents · AI governance & market-conduct evidenceAGT-INS-16 · Agent Care available after launch