Nestack Agent Care
Industries / Human Resources / Candidate screening agent

HR AI agent · Candidate screening

Candidate Screening AI Agent

Score applications against the written requisition, keep the ranking, the notice and the impact figures together, and hand the shortlist to the named recruiter who decides who is interviewed.

4–6 weeksTypical delivery
Your stackDeployment
Notice firstNamed recruiter
Agent CareAfter launch

What this agent does

Produces the rank, not the rejection

In
01

The written requisition sets the criteria, and the score records which of them moved it.

02

The output is a rank, and where the line falls under that rank is drawn by a person.

Reason
03

Selection rate, scoring rate and impact ratio go to the record, and so does the unknown count.

04

A fit score makes the tool an automated employment decision tool under NYC Admin. Code 20-870.

05

Candidate notice runs ten business days ahead of use, so a requisition without one is not scored.

Decide
06

A vendor-supplied score is treated as a consumer-report question, not as a product feature.

07

Postcode, school and gaps in work history are flagged as proxies before they are weighed.

Out
08

A ranking carries the day it ran and the adverse-impact benchmark in force on that day.

09

Write to the tracker only inside the approval boundaries agreed during implementation.

Product statement

Reading, scoring and record-keeping belong to the agent. The shortlist belongs to a named recruiter, who takes it into the hiring decision and owns it there.

Example workflow

One requisition, application to shortlist

AgentHuman
1Applications receivedCVs, application forms, screening answers, referral notes and tracker fields
2Criteria fixed and versionedThe written criteria, the weights, the notice already served and the day that set was fixed
3Scored and ranked apartThe score, the rank, the features that moved it and the criteria it was read against
4Controls appliedProxy checks, notice checks, impact-ratio checks and scoring confidence
No human action required

Stages 1 to 4 run unaided, and nobody leaves the pool at any of them — the agent is ranking, and the recruiter lane opens at the shortlist gate.

5DecisionSplits at the shortlist gate
Impact ratios within tolerance

Goes to the named recruiter to shortlist.

Anything skewed

Adds a talent-lead read first.

Recruiter review

The ranking is held with its criteria, its impact figures and the applications it was drawn from.

Shortlist · Return to pool · Send to talent lead
Shortlisted — by the named recruiter
6Tracker and notice records updatedOnly where write access and hiring records policy allow it
7Outcome evaluatedImpact ratios, notice timing, recruiter corrections and what review found
Corrections

Each recruiter correction is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Rejecting an applicant on a score.
Deciding who is interviewed.
Setting where the shortlist line falls.
Signing a bias audit or its published summary.
Automation boundaryAgent acts unaided
Read each application against the written criteria.
Record the selection rate and the scoring rate for each category.
Hold the ranking until the notice is evidenced.
Flag postcode, school and gaps in work history as possible proxies.
Nobody leaves the pool except by a named recruiter, inside the agreed boundaries.
Judging whether a criterion is job related.
Telling a candidate why they did not advance.
Choosing which requisitions the agent scores.
Changes to the criteria, weights or notice.

Example output

One scored application, annotated

This serves a talent team who may have to explain a rejection years after the requisition closed; below is one scored application exactly as the agent leaves it.

Scored application · single requisitionIllustrative example
Score
Recorded as
Requisition
Evidence of record
Confidence
Held for
Data engineer, London office
Scored above the sample median
A rank, not a decision
Application form, 3 August 2026
Held unshortlisted
The named recruiter, by name
As receivedBuilt from the application form and the written requisition alone, and it claims nothing beyond them.
What the record holds Application form Written criteria Notice served
Why no rejection hereCutting a candidate from the pool is a judgement the recruiter makes.
ActionShortlistReturn to poolSend to talent lead
What the score decidesBelow the set threshold a ranking gets a talent-lead read before the recruiter sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Each scoreFrom the application that carries it
03Evidence

Where the evidence is used

The agent does not vouch for an applicant, only for what the application says, when it arrived, and how the score was built from it.

01Approved path

It learns who got hired

Train a model on a decade of past hires and it reproduces them fluently: change nothing on a CV but the name, and the rank moves.

02Human review

What was checked, and not found

The four-fifths rule sits at 29 CFR 1607.4(D) and was read there on 20 August 2026, but OFCCP removes its own copy at 41 CFR 60-3 on 26 October 2026 and the EEOC has a rescission at final-rule stage, so a ranking records which benchmark it was measured against and the date that benchmark applied.

04Build an evidence trail

The application, the score it drew and the recruiter who shortlisted it stay together.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Applicant trackingGreenhouse · Lever · Workday
Applications, stages and outcomes
Requisition and criteriaJob specs · rubrics · weights
Written criteria of record
Candidate-supplied documentsCVs · forms · screening answers
What the applicant chose to send us

Agent

Candidate screening and ranking

Reads the applications
Scores and ranks
Holds for the recruiter

Vendor scoring servicesAssessment vendors · scoring APIs
Third-party scores and summaries
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six sieves between the model and the recruiter

Six sieves down one column, the last the finest. What settles is drawn in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeReturn the pool unranked when evaluation or production signals degrade.Roll back
L5Version monitoringTrack model, prompt and criteria rules, and note the version each score was produced under.Track
L4TraceabilityRecord each score, the application under it, the criteria it used and every read of the file.Record
L3Recruiter releaseHold the ranking for a named recruiter; the hold governs release, not whether the rank is right.Gate
L2Proxy guardrailsTest each score against its versioned criteria, and refuse a rank that moved on a flagged proxy.Restrict
L1Impact thresholdsRoute a skewed impact ratio to a talent-lead read before the ranking reaches a shortlist.Require review
Model coreRanking produced — the criteria, the scores, the impact figures and the notice record
L1 – L2Test whether a ranking may stand
L3Leaves the shortlist to a named recruiter
L4 – L5Keep the shortlist and the score behind it
L6Ranks nothing and returns the pool when signals degrade

How Nestack evaluates it

Evaluate the whole screen — not only the shortlist that comes out.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the shortlist a recruiter works
Depth of coverage ▼
E1Final-output evaluationDid the ranking record the criteria it was actually built on?
E2Step-level evaluationDid the agent read the right requisition, the right criteria and the live weights?
E3Tool evaluationDid it read and write the correct requisition and the correct application?
E4Confidence calibrationDo low-confidence scores actually attract more recruiter corrections?
E5Slice evaluationHow does performance change across specific requisition classes?
E6Business outcomeHow many rankings needed a correction before the recruiter shortlisted?
Floor — the pool a ranking rests on

Failure modes

Where each failure originates in the agent

Seven failure modes, set at the stage where each one first shows itself.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
TH-03

Scored before notice served

The application is read with no notice on file.

Stage gathersThe forms, the criteria and the notice record
02 · Reasoning2 modes
TH-04

Proxy carries what the score drops

A postcode or a gap in work stands in for a person.

TH-06

Rescinded benchmark read as live

A withdrawn measure is worked as the current one.

Stage proposesThe criteria, the score and the impact figures
03 · Tool / write2 modes
TH-02

Thin impact base passed forward

A ranking moves on with most categories unknown.

TH-05

Vendor score acted on unflagged

A bought-in score reaches a rejection unmarked.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
TH-01

Shortlisted, ranking unrecorded

The record shows a shortlist but not the ranks under it.

Stage returnsThe ranking a recruiter reads and the notice sent
05 · Change / Version1 mode
TH-07

Silent drift into a scored tool

A summariser gains a score and changes what it is.

Stage tracksModel, prompt, criteria rules and score dates
Sev-1 · a candidate cut on no record Sev-2 · a skewed ranking reaches a lead Sev-3 · signals degrade, pool held unranked

Affected slices

Entry-level volume absorbs the corrections

A requisition-level impact figure can read clean while high-volume entry-level roles carry most of the corrections. Nestack reports the correction rate by requisition class, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
High-volume entry-level roles14.3%3.7× Review
Multi-jurisdiction requisitions10.1%2.6× Review
Career-changer applications6.3%1.6× Watch
Established specialist roles2.9%0.8× Normal
Bar: correction-rate lift vs. specialist-role baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What a thinned pool costs

A cycle ends when the group thinned at one stage alone is a standing case. That suite is what the next requisition run is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Correction rate rises on high-volume entry-level roles.

02Diagnose

The requisition whose shortlist quietly stopped holding anyone with a break in their work history is worked backwards until one cause is left standing.

03Improve

The shortlist goes out numbered, and the scores beneath it ride with it.

04Verify

One impact case still failing is enough to hold the shortlist back.

05Learn

It is kept for good, and the scoring rules change in that same commit.

Learn → DetectThe return edge. The next ranking is measured against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, scoring and ranking, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Requisition criteria and automation-boundary scope.
02Tracker, requisition and vendor sources.
03Application-to-score and criteria-coverage mapping.
04Applicant source ingestion.
05Criteria, weight and score binding.
06Impact scoring and review routing.
07Recruiter shortlist workflow.
08Applicant-tracker integration.
09Impact and scoring cases.
10Guardrails and shortlist controls.
11Score-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne requisition, one cycle ProductionProduction hiring workflow AdvancedMultiple requisitions / jurisdictions
Introduced at Pilot
Scoring to your written criteria
Named recruiter shortlisting
Applicant-pool baseline
Introduced at Production
Reporting by requisition class
Recruiter review workflow in your systems
Approved write-back
Tracker-and-source integration
Introduced at Advanced
Multi-source applications
Cross-requisition ranking packs
Large applicant pools
Multi-jurisdiction notice controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, application volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your live requisitions and the criteria each is scored against Criteria capture and notice versioningWeek 1
02Representative applications, tracker fields and vendor scores Source binding, scoring logic and the applicant-pool baselineWeek 2
03Your hiring calendar and the recruiters it names Criteria mapping, notice binding and the automation boundaryWeek 1
04Access to the tracker APIs, feeds or exports Tracker, requisition and vendor source assessment, then integration setupWeek 2
05Rankings you would not want audited Impact cases and the failure roundWeek 4
06What no score may establish Impact scoring, review routing, guardrails and release controlsWeek 3
07A named recruiter who shortlists Release to the named recruiter, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

The fifth week carries two bands because those two phases coincide, not to make the column read better.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Criteria discovery, notice versioning and the automation boundary W2Source integration and the applicant-pool baseline W3Scoring, ranking logic and release controls W4Evaluation suite, impact cases and the failure round W5Tracker integration, pilot rankings and targeted corrections W6One hiring cycle run under the talent lead, then Agent Care handover
Reading the bandNo bar runs past the weeks its own work is named for, and week five is shared by design.
At the end of W6When the score record validates, Agent Care adopts the agent.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · HR AI agent

Build a candidate screening agent around the ranking your last requisition never wrote down.

Show us one requisition you run each quarter and the criteria behind the last shortlist. If nobody can name the benchmark it was measured against, the answer is a date rather than a rule: Title VII has not moved, and the instrument that measured it is leaving.

Nestack Agents · Candidate screeningAGT-HR-01 · Agent Care available after launch