Nestack Agent Care
Industries / Travel & Hospitality / Testing programme agent

Travel & Hospitality AI agent · Aviation testing

Aviation Testing Programme AI Agent

Tally every safety-sensitive name into the pool with the date it entered or left, tie each test result to the pool it came from, and hold the report for the Designated Employer Representative.

4–6 weeksTypical delivery
Your stackDeployment
March deadlineNamed DER
Agent CareAfter launch

What this agent does

Assembles the report, never the certification

In
01

A test is ordered, and § 120.105 sets which duties put an employee in the pool.

02

A reporting year closes, and § 120.119 puts the MIS report with the FAA by 15 March.

Reason
03

A report is prepared, and § 40.26 names the form at appendix J to 49 CFR part 40.

04

A year of results is certified, and § 120.109(b) and § 120.217(c) set next year off it.

05

A rate notice publishes, and 90 FR 56249 of 5 December 2025 governs the current year.

Decide
06

A collection is ordered on oral fluid, and two HHS-certified laboratories are required first.

07

An oral fluid test cannot run, and 91 FR 25507 of 11 May 2026 finds no laboratory certified.

Out
08

A return-to-duty case opens, and 91 FR 10518 of 4 March 2026 is a notification, not a rule.

09

Execute write actions only inside the approval boundaries agreed during implementation.

Product statement

The agent holds the pool census and reconciles the tests to it. The Designated Employer Representative certifies the accuracy and completeness of the MIS report.

Example workflow

One pool month, evidence to certification

AgentHuman
1Pool and test evidence receivedRoster feeds, selection lists, collection records or laboratory results
2Pool context assembledThe employee, the duty performed, the date the name entered the pool and the date it left
3Report evidence draftedThe tests reconciled, the pool months behind them, the gaps and completeness
4Controls appliedCensus checks, reconciliation checks, reporting-clock checks and completeness confidence
No human action required

Stages 1 to 4 run unaided, and nothing is certified at any of them — the agent is assembling, and the compliance lane opens at the completeness gate.

5DecisionSplits at the completeness gate
Evidence sufficient

Goes to the Designated Employer Representative to certify.

Anything thin

Adds a compliance analyst read first.

Compliance review

The report is held with its pool months, its unreconciled tests and the categories they fall in.

Certify · Append evidence · Send to compliance
Certified — by the named representative
6Pool and test records updatedOnly where write access and records policy allow it
7Outcome evaluatedCensus accuracy, test reconciliation, analyst corrections and what the read found
Corrections

Each compliance correction is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Certifying the MIS report under § 120.119.
Deciding that a duty is not safety-sensitive.
Submitting the annual report to the FAA.
Removing a name from the testing pool.
Automation boundaryAgent acts unaided
Hold the pool census with the date every name entered.
Reconcile every test result to the pool month that produced it.
Track the reporting clock against the 15 March date that closes it.
Flag the month whose pool census is not evidenced.
Nothing is certified or submitted except by a named person, inside the agreed boundaries.
Judging whether the report may be certified.
Telling the FAA a testing programme is in order.
Setting the rate a testing programme must meet.
Changes to pool, roster or test records.

Example output

One pool month, annotated

The reporting deadline falls on 15 March; this record is what a single pool month carried into the report.

Report evidence · single pool monthIllustrative example
Pool month
Recorded as
Test category
Evidence of record
Confidence
Held for
Random selection, § 120.109
Reconciled to the pool, census evidenced
Random drug
Collection record, 3 August 2026
Held uncertified
The Designated Employer Representative
As receivedTaken from the roster feed and the collection record — it reaches as far as those sources do.
What the record holds Pool census date Collection record Laboratory result
Why no certification hereWhether the report may be certified is a § 120.119 act, not a model output.
ActionCertifyAppend evidenceSend to compliance
What the score decidesBelow the configured threshold the report picks up a compliance read before the DER sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every testFrom the pool that produced it
03Evidence

Where the evidence is used

Our driver qualification and Clearinghouse agent runs per-driver queries on each driver's own twelve-month clock; this page is a per-company annual certified aggregate that sets the following year's random rate, a feedback loop with no FMCSA analogue.

01Approved path

This year sets next year

What the representative certifies in March is what § 120.109(b) and § 120.217(c) work from: 90 FR 56249 of 5 December 2025 was signed by Brett A. Wyrick, Deputy Federal Air Surgeon.

02Human review

What was checked, and not found

Retention under § 120.219 was not retrieved; the certification block on appendix J itself was not read; the status of the October 2024 electronic-signatures NPRM is unknown; and a laboratory certified after 11 May 2026 could not be ruled out.

04Build an evidence trail

The test, the pool it was drawn from and the representative who certified stay on file.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Roster and duty recordsHR system · duty rosters
Safety-sensitive duty codes
Collection and laboratoryCollection sites · custody forms
Laboratory results
Random selectionConsortium · C/TPA feeds
Selection lists and draws

Agent

Aviation testing programme

Reads the pool
Reconciles the tests
Holds for the DER

Reporting and returnsMIS submissions · exports
Operations specification
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six audits between the model and the representative

Six audits of the same pool, and no two are alike. Whatever counts is set out in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeNarrow the agent to pool assembly when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt and pool rules; the fentanyl panel expansion is an NPRM of 2 September 2025 and is not in force.Track
L4TraceabilityRecord each test, the pool it came out of, the month it falls in and every read of the report.Record
L3DER releaseHold the report for the Designated Employer Representative; the hold governs release, not accuracy.Gate
L2Scope guardrailsTest the evidence against 14 CFR part 120 and 49 CFR part 40; the 2026 part 40 rule was signed by Sean P. Duffy.Restrict
L1Confidence thresholdsRoute a thin report to a compliance read; oral fluid is authorised in the rule and unusable in practice, with no HHS-certified laboratory at 91 FR 25507, 11 May 2026.Require review
Model coreEvidence assembled — the pool months, the tests, the categories and completeness
L1 – L2Test whether a report may stand
L3Puts the certification in a person's hands
L4 – L5Keep the test and the pool behind it
L6Stops at pool assembly when signals degrade

How Nestack evaluates it

Evaluate the whole assembly — not only the report evidence that comes out.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the report the FAA reads
Depth of coverage ▼
E1Final-output evaluationDid the evidence record what each pool month actually held?
E2Step-level evaluationDid the agent read the right roster, the right month and the live selection list?
E3Tool evaluationDid it read and write the correct pool and the correct test result?
E4Confidence calibrationDo low-confidence reports actually attract more compliance corrections?
E5Slice evaluationHow does performance change across specific test categories?
E6Business outcomeHow many reports needed a correction before the representative certified?
Floor — the report the employer answers for

Failure modes

Where each failure originates in the agent

Seven failure modes, each set where it first shows in the run.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
KC-03

Stale roster read

The roster read is not the pool now in force.

Stage gathersThe people, the tests, the pools and the dates
02 · Reasoning2 modes
KC-04

Report asserted, not shown

A month is called evidenced without its census.

KC-06

Proposed panel read as live

The fentanyl NPRM is worked as though final.

Stage proposesThe tests, their pools and completeness
03 · Tool / write2 modes
KC-02

Thin report passed forward

A report moves on without the compliance read.

KC-05

Test bound to wrong pool

A result is filed against the wrong month.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
KC-01

Certified, evidence unrecorded

The report shows a certification but not its basis.

Stage returnsThe report a representative certifies and the FAA reads
05 · Change / Version1 mode
KC-07

Silent census regression

A configuration change moves the pool, not the report.

Stage tracksModel, prompt, pool rules and report fields
Sev-1 · a report certified on no census Sev-2 · wrong test reaches the report Sev-3 · source degrades, report uncertified

Affected slices

Random selections absorb the corrections

A pool-level census-accuracy figure can read clean while random selections carry most of the rework. Nestack reports the correction rate by test category, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Random drug selections7.7%3.6× Review
Pre-employment tests5.5%2.6× Review
Post-accident tests3.5%1.6× Watch
Follow-up and return-to-duty1.4%0.7× Normal
Bar: correction-rate lift vs. follow-up baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What a mis-sized pool costs

A cycle closes when the mis-sized pool is a regression case. That suite is what the next report assembled is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Correction rate rises on random selections.

02Diagnose

The headcount submitted in March that set a testing rate for the whole of the next year is read back until one cause remains.

03Improve

The change ships numbered, and the tests that forced it ride with it.

04Verify

Nothing releases while one touched pool case is still red.

05Learn

It is retained for good, and the pool rules are amended in that same commit.

Learn → DetectThe return edge. The next testing year is measured against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, report assembly, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Safety-sensitive pool discovery and boundary work.
02Roster, collection and result sources.
03Test-to-pool and reporting-clock coverage mapping.
04Test and pool-census ingestion.
05Employee, duty and pool binding.
06Completeness scoring and review routing.
07Representative certification workflow.
08Roster and testing-system integration.
09Pool-size and rate cases.
10Guardrails and representative controls.
11Test-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne pool, one testing year ProductionProduction reporting workflow AdvancedMultiple pools / employers
Introduced at Pilot
Report assembly to your pools
Representative release
Safety-sensitive pool baseline
Introduced at Production
Reporting by test category
Certification workflow in your systems
Approved write-back
Collection-site integration
Introduced at Advanced
Multi-employer consortia
Cross-year evidence packs
Large safety-sensitive rosters
Multi-employer pool controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, pool size, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your safety-sensitive employees and the duties they hold Pool inventory mapping and census captureWeek 1
02Representative roster, collection and result records Test binding, pool logic and the census baselineWeek 2
03Your duty list under § 120.105 and § 120.215 Pool mapping, test binding and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Roster, collection and result-source assessment, then integration setupWeek 2
05Pools you would not want counted Rate cases and the evaluation runWeek 4
06What no testing report may establish Completeness scoring, review routing, guardrails and release controlsWeek 3
07A Designated Employer Representative to certify Certification workflow, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Each span is a week the work really needs and not a drawing, so one week ends up carrying a pair.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Testing programme discovery, pool mapping and the automation boundary W2Source integration and the census-accuracy baseline W3Report assembly, reconciliation logic and release controls W4Evaluation suite, category cases and failure-mode testing W5Record integration, pilot pools and targeted corrections W6One testing year run under the designated representative, then Agent Care handover
Reading the bandEach bar covers only the weeks its own work is named for. The fifth week carries two because the work does.
At the end of W6Validation closes on live pools, and Agent Care picks up the watch.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Travel & Hospitality AI agent

Build a testing programme agent around the report your representative has to certify.

Show us one pool month and the tests drawn from it. Send us the census behind your last March filing and we will show you where the evidence for it stops. Oral fluid is authorised in the regulation and unusable in practice; hazard registers and safety policy are another agent.

Nestack Agents · Aviation testingAGT-TH-20 · Agent Care available after launch