Nestack Agent Care
Industries / Quality Assurance / Visual inspection and AOI

Quality assurance AI agent · Inspection capability

Visual Inspection and AOI AI Agent

Treat the model as an instrument: record what it called, under which version, measure agreement class by class, and leave the validation to the process owner whose name goes on it.

4–6 weeksTypical delivery
Your stackDeployment
Gauge firstNamed owner
Agent CareAfter launch

What this agent does

Measures the instrument, never validates it

In
01

A pass or fail call is attribute data, produced by an instrument, not an opinion.

02

No reference artefact exists for an acceptable joint, so agreement is the only truth there is.

Reason
03

Section 820.75 was reserved on 2 February 2026, and the duty moved into ISO 13485 clause 7.5.6.

04

FDA quoted clause 7.6 at Linemaster on 27 May 2026: software used for monitoring and measurement.

05

Computer Software Assurance went final on 24 September 2025 and superseded Section 6 of the GPSV.

Decide
06

Annex III omits quality control; inspection is a quality function, not a safety function.

07

Regulation (EU) 2026/1744 moved machinery to Section B and deferred Annex III to 2 December 2027.

Out
08

No published method joins classifier scoring to measurement-system analysis, and that is the finding.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

Assembly, agreement measurement and versioning belong to the agent. The validation, and the name that goes under it, belong to the process owner an investigator later asks for.

Example workflow

One instrument, images to validation

AgentHuman
1Evidence receivedLabelled images, the reference-truth method, the model version, the optics record and the station
2Agreement study assembledThe held-out set, a temporally later slice, the class balance present and who resolved disagreements
3Capability evidence raisedThe confusion matrix, escapes and false calls, read per defect class and never as one figure
4Controls appliedBalance checks, domain-shift checks, version checks and evidence completeness
No human action required

Stages 1 to 4 run unaided, and nothing is validated at any of them — the agent is measuring, and the owner lane opens at the validation gate.

5DecisionSplits at the validation gate
Agreement inside the criteria

Goes to the process owner to validate.

Anything under-sampled

Adds a quality-engineering read first.

Process-owner review

The pack is held with its images, its per-class results and the studies the agent could not run.

Validate · Extend evidence · Send to engineering review
Validated — by the named process owner
6Quality-system and model records updatedOnly where write access and records policy allow it
7Outcome evaluatedEscapes found downstream, false calls on the station, owner corrections and what the later review turned up
Corrections

Each correction the owner makes counts in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Signing the validation of an inspection process.
Dispositioning a part on the strength of a call.
Setting the agreement criteria a study must meet.
Declaring an instrument capable for its intended use.
Automation boundaryAgent acts unaided
Record the model version each call was made under.
Run the agreement study against an independently set reference.
Report escapes and false calls for each defect classification separately.
Hold the evidence pack for the named process owner.
Nothing is validated for production except by a named owner, inside the agreed boundaries.
Judging whether a learned classifier is a gauge.
Telling an investigator the record is complete.
Choosing which defect classes the catalogue admits.
Changes to the validation rules or the release gate.

Example output

One capability pack, annotated

This serves a quality team who may have to defend an instrument to an investigator years after it was validated, against a manual that is industry practice rather than law; below is one pack exactly as the agent leaves it.

Capability pack · single instrumentIllustrative example
Instrument
Recorded as
Defect class
Evidence of record
Confidence
Held for
Surface inspection, one station
Agreement measured, not proven
Per class, not overall
Held-out slice, 6 August 2026
Held unvalidated
The process owner, by name
As receivedDrawn from the labelled set, the held-out slice and the optics record, and it claims nothing past them.
What the record holds Labelled image set Held-out slice Optics configuration
Why no validation hereCalling an instrument capable is a judgement the process owner makes.
ActionValidateExtend evidenceSend to engineering review
What the score decidesBelow the configured threshold a pack gets an engineering read before the owner sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Each callFrom the instrument that made it
03Evidence

Where the evidence is read

Unomedical was cited on 8 January 2026 under 820.75(a) for an inspection test method never validated, against more than five thousand complaints — three weeks before that section ceased to exist.

01Approved path

A model is a gauge

Its own work instruction required that test to “consistently and accurately differentiate between the possible test states”, which is what a gauge is.

02Human review

What was checked, and not found

ISO/IEC TS 4213 scores classifiers and does not touch measurement-system analysis; AIAG MSA-4 is 2010 and predates deep learning; IEEE P2975.2, which would join them, has held an active PAR since 15 February 2023 and is still unpublished. No independent false-call rate was found.

04Build an evidence trail

The image, the call the model made and the inspector who confirmed it stay together.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Inspection stationsAOI · surface · assembly cells
Captured images and their calls
Labelled evidenceImage sets · label stores
Reference truth, and who set it
Quality system recordseQMS · validation protocol files
Validation records and the triggers

Agent

Visual inspection capability

Reads the instrument
Assembles the evidence
Holds the validation

Model and training registryModel registry · data versions
Model versions and the training cuts
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six lenses between the model and the station

Six lenses on one bench, the last the sharpest. What resolves is set out in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeNarrow the agent to evidence assembly when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt and training-set versions, and note the version each call was made under.Track
L4TraceabilityRecord each call, the image under it, the model version and every read of the file.Record
L3Owner releaseHold the pack for a named process owner; the hold governs validation, not whether the model is capable.Gate
L2Agreement guardrailsTest each pack against its agreement criteria, and return one whose defect class is too thin to measure.Restrict
L1Confidence thresholdsRoute a thin or shifted evidence pack to an engineering read before the owner sees it.Require review
Model corePack assembled — the images, the per-class results, the versions and what stays unmeasured
L1 – L2Test whether an instrument may stand
L3Leaves the validation to a named owner
L4 – L5Keep the call and the image behind it
L6Sends the part to a human bench when signals degrade

How Nestack evaluates it

Evaluate the whole capability run — not only the call that comes out.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the call a station acts on
Depth of coverage ▼
E1Final-output evaluationDid the pack stay inside what the labelled images actually established?
E2Step-level evaluationDid the agent read the right model version, the right images and the live criteria?
E3Tool evaluationDid it read and write the correct call and the correct validation record?
E4Confidence calibrationDo low-confidence calls actually attract more owner corrections?
E5Slice evaluationHow far does agreement fall when the defect class changes?
E6Business outcomeHow many packs needed a correction before the owner validated one?
Floor — the pack read back at audit

Failure modes

Where each failure originates in the agent

Seven failure modes, each set where the bench first reveals it.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
VI-03

Truth set by the same eyes

The study compares the model against itself.

Stage gathersThe images, the labels, the optics and the station
02 · Reasoning2 modes
VI-04

Accuracy read off a thin class

A model that passes parts still scores well.

VI-06

Reserved section worked as live

A retired rule number is cited as current.

Stage proposesThe call, the class, the version and the evidence
03 · Tool / write2 modes
VI-02

Offline score not seen on the line

Performance decays on later production.

VI-05

Unlabelled defect class passes

A defect nobody catalogued reads as normal.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
VI-01

A light moved, nobody looked

Optics change and the calls change with them.

Stage returnsThe call a station acts on and an owner signs
05 · Change / Version1 mode
VI-07

Silent revalidation trigger

A model version ships outside change control.

Stage tracksModel, prompt, training cuts and criteria
Sev-1 · a class nobody labelled ships Sev-2 · a light moves, the calls shift Sev-3 · agreement thin, pack held back

Affected slices

Unseen classes absorb the corrections

A station-level agreement figure can read clean while unseen defect classes carry most of the owner corrections. Nestack reports the correction rate per defect class, and not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Unseen defect classes8.5%3.7× Review
Changed lighting or optics6.0%2.6× Review
New supplier material3.7%1.6× Watch
Catalogued defect classes1.8%0.8× Normal
Bar: correction-rate lift vs. catalogued-class baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What an unmeasured class costs

A cycle shuts when the defect class nobody ever labelled is a standing case. That suite is what the next validation run is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Correction rate rises on unseen defect classes.

02Diagnose

The instrument that passed parts for a month because a lamp had aged, and read confident throughout, is worked backwards until one cause is left standing.

03Improve

Runs ship numbered, and the images behind them travel attached.

04Verify

Each touched agreement case is run again, and one red holds it back.

05Learn

The case is kept, and the validation rules change in that same commit.

Learn → DetectThe return edge. The next validation run meets a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, evidence assembly, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Model-as-instrument and automation-boundary scoping.
02Image, label and model-version sources.
03Clause-to-evidence and revalidation-trigger mapping.
04Image and label intake.
05Class binding and pack logic.
06Agreement scoring and review routing.
07Process-owner validation workflow.
08Quality-system integration.
09Agreement and drift cases.
10Guardrails and disposition controls.
11Call-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne defect class, one cycle ProductionProduction inspection workflow AdvancedMultiple stations / lines
Introduced at Pilot
Evidence assembly to your criteria
Named process-owner validation
Defect-catalogue baseline
Introduced at Production
Reporting by inspection class
Engineering review workflow in your systems
Approved quality-record write-back
Camera-and-line integration
Introduced at Advanced
Multi-regime validation rules
Cross-station evidence packs
High image volume
Multi-line validation controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, image volume, validation controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your live stations and the defect classes each one calls Criteria capture and model versioningWeek 1
02Representative images, labels and past validation records Class binding, pack logic and the defect-catalogue baselineWeek 2
03Your change-control path and the process owner it names Criteria mapping, class binding and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Station, label and quality-record source assessment, then integration setupWeek 2
05Calls you would not want re-read Agreement cases and the evaluation roundWeek 4
06What no call may establish Agreement scoring, review routing, guardrails and release controlsWeek 3
07A named process owner who validates the instrument Handover to the process owner, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

A band is exactly as wide as its phase costs here, so week five shows a pair where one would look neater.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Criteria discovery, model versioning and the automation boundary W2Station and label integration and the defect-catalogue baseline W3Class binding, pack logic and release controls W4Evaluation suite, agreement cases and failure-mode testing W5Quality-system integration, pilot packs and targeted corrections W6One validation cycle run under the process owner, then Agent Care handover
Reading the bandWeek five is the only place two bars sit together, and no bar is stretched to fill a row.
At the end of W6When the validation record validates, Agent Care picks the agent up.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Quality assurance AI agent

Build a visual inspection agent around the validation your last model was never given.

Show us one inspection a model already decides and the record behind it. Not how it scored on a benchmark. Which defect classes sat in the labelled set, who set the reference truth, and whose name is on the validation. A class nobody labelled comes back as a case.

Nestack Agents · Inspection capabilityAGT-QA-04 · Agent Care available after launch