Assemble the evidence behind an exam sitting, check it against the accommodations on file, and hold each flag for the integrity officer, who decides whether anything is a charge.
The first four stages run unaided, and no charge exists at any of them — the agent is describing a sitting, and the officer's lane opens at the confidence gate.
5DecisionBranches at the flag threshold
Weak signal
Goes to the integrity officer to review.
Strong signal
Adds a disability-services check first.
Officer review
The sitting is held with its timestamped windows and the behaviour described in words.
Dismiss · Annotate · Send to disability services
Reviewed — recorded against the sitting▼
6Review record updatedOnly where write access and review policy allow it
7Sitting evaluatedDismissal rate, cohort flag rates, accommodation misses and post-review corrections
Dismissals
Each reviewer dismissal is counted in the evaluation.
What should not run autonomously
Human approval stays in control
Outside the boundary — human approval required8 items
Charging a student with academic misconduct.
Assigning a grade, a penalty or a transcript notation.
Requiring, viewing or storing a room scan.
Granting, denying or changing a testing accommodation.
Automation boundaryAgent acts unaided
✓Assemble the sitting's artefacts and index them by timestamp.
✓Suppress the signal classes an approved accommodation.
✓Describe observed behaviour in words.
✓Hold the flagged sitting for the integrity officer, record intact.
Any write happens inside the boundaries agreed at implementation, never ahead of review.
Ending a sitting, or refusing entry to one.
Identifying a person from a face, voice or biometric template.
Deciding an appeal, or the outcome of a hearing.
Changing a detection threshold inside a live exam window.
Example output
One flagged sitting, annotated
Everything the agent reports is attached to the artefact it was drawn from.
Integrity output · single sittingIllustrative example
Sitting
Observed behaviour
Evidence window
Source artefact
Confidence
Accommodation check
Proctored final
Looked away from the screen repeatedly through the second half
00:18:42
Recorded session
88%
No suppression on file
As receivedTaken from the recorded session and the accommodations record.
Evidence usedRecorded session windowStated sitting conditionsAccommodations record
Why this flagIt describes what the artefact shows, not what the student intended.
ActionDismissAnnotateSend to disability services
What the score decidesBelow the configured threshold the sitting picks up a disability-services check before.
Value
Where AI adds value
The same four claims, placed at the point in the workflow where each one applies.
Where the value landsValue 01 – 04
Every sittingFrom the assessment platform
03Observation
Describe from the artefact
Work from the recorded session, the stated sitting conditions and the accommodations on file.
01Approved path
Flag the sitting, not the student
Routine sittings clear without anyone watching a recording end to end.
02Human review
Send review where risk concentrates
Flagged windows and low-confidence sittings are marked, so the officer's read starts at the evidence.
04Build an evidence trail
The flag keeps its evidence window, its confidence and its reviewer.
Integrations
Typical integrations
Five system groups connect to the same agent. Which of them are in scope is decided in discovery.
An aggregate flag rate can look defensible while a few cohorts carry most of the flags Nestack reports it by slice rather than in aggregate The cohorts that carry it are named, not averaged away..
Slice performance — reported separately, not only in aggregateIllustrative example
Slice
Failure rate
Lift
Lift vs. threshold
Status
Students using assistive technology
9.2%
2.8×
Review
Students the face model reads worst
6.9%
2.1×
Review
Non-native English writers
5.9%
1.8×
Watch
Students in none of these cohorts
2.3%
0.7×
Normal
Bar: flag-rate lift vs. the all-sitting baseline · scale 0–4.0× · tick marks the 2.0× review threshold2 of 4 slices over threshold
Evidence-linked improvement
Every cycle leaves a case behind
An explanation closes nothing here. The cycle ends with a case the next release has to clear.
Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect
Flag rate rises in one student cohort.
02Diagnose
The recording, the suppression log and the policy version are read side by side until one of them explains it.
03Improve
The change is stamped to a version, with the flagged sittings attached.
04Verify
A failing case stops the release, not a reviewer's judgement.
05Learn
The case sticks, and the flag-review standard moves with it.
Learn → DetectThe return edge. Every later detection is measured against the larger suite.
Typical build scope
Twelve workstreams across six weeks
The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, artefacts, flagging workflow, evaluation, integration, then production validation and handover.
WorkstreamWeek 1Week 2Week 3Week 4Week 5Week 6
01Integrity workflow discovery and boundary definition.
02Assessment and proctoring source review.
03Integrity policy and accommodation-suppression mapping.
04Sitting-artefact ingestion and indexing.
05Observation logic and evidence windows.
06Confidence scoring and flag routing.
07Integrity officer review workflow.
08Assessment-platform integration.
09Flag regression cases.
10Guardrails and charge controls.
11Evidence-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallelFinal scope and sequence confirmed in discovery
Engagement tiers
What each tier includes
Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.
Capability✓ in scope · — not at this tierPilotOne exam window, one schoolProductionProduction assessment systemsAdvancedMultiple campuses / systems
Introduced at Pilot
Flagging to your policy and record✓✓✓
Officer review✓✓✓
Flag-quality baseline✓✓✓
Introduced at Production
Reporting by sitting type—✓✓
Review workflow in your systems—✓✓
Approved write-back—✓✓
Assessment-platform integration—✓✓
Introduced at Advanced
Multi-campus and multi-state rules——✓
Multi-stage conduct review——✓
High sitting volume——✓
Enterprise integrity controls——✓
Build priceFrom $5,000From $8,000Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, sitting volume, review controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.
What we need from you
What you bring, and what we build with it
Each input maps to a piece of build scope and a week in the delivery timeline.
You bringWe build with it
01Your integrity policy and its sanction ladder→Integrity policy and suppression mappingWeek 1
02Representative proctored sittings→Flagging baseline, signal extraction and evidence-window bindingWeek 2
03Your accommodations record and its update path→Integrity policy and accommodation-suppression mappingWeek 1
04Access to relevant APIs, feeds or exports→Proctoring, assessment and roster assessment, then integration setupWeek 2
05Flags that were not misconduct→Paired-sitting cases and the evaluation suiteWeek 4
06What a flag must reach before it becomes a charge→Confidence scoring, flag routing, guardrails and review controlsWeek 3
07Named integrity officers to review flags→Officer review workflow, then pilot and production validationWeeks 5–6
Nothing else is requiredDeployment, documentation and Agent Care handover are ours.
Delivery timeline
Four phases across six weeks
Bands track the weeks work actually runs in, so the fifth carries two phases rather than one padded one.
PhaseW1W2W3W4W5W6
DiscoveryW1
BuildW2 – W3
EvaluateW4 – W5
Pilot & LaunchW5 – W6
Week focusW1Integrity workflow discovery, policy mapping and the automation boundaryW2Artefact integration and the flagging baselineW3Flagging workflow, confidence logic and review controlsW4Evaluation suite, cohort flag-rate checks and failure-mode testingW5Platform integration, a pilot exam window and targeted correctionsW6One exam window reviewed by the integrity officer, then handover
Reading the bandEach band covers only the weeks its work is named in. The doubled fifth week is real, not padding.
At the end of W6Validation complete; Agent Care holds the monitoring from week seven.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.
Next step · Education AI agent
Build a proctoring agent around your institution's conduct process.
Show us your assessment platform, your integrity policy and who reads a flag. Nothing reaches a conduct file until we have mapped the review path, set the automation boundary and named what stays with the officer.