Nestack Agent Care
Industries / Human Resources / Compensation benchmarking agent

HR AI agent · Compensation benchmarking

Compensation Benchmarking AI Agent

Draw the range, name the survey it came from and count the employers standing behind the cut, then hold that number, with its provenance attached, for a named reward officer to approve.

4–6 weeksTypical delivery
Your stackDeployment
Source tracedNamed officer
Agent CareAfter launch

What this agent does

Draws the range, never approves it

In
01

A range is drawn, and the survey beneath it is named on the face of the output.

02

A source cannot be traced to a survey, and it is marked unusable rather than blended in.

Reason
03

A benchmark thins to a few reporters, and that reporter count is printed beside the number.

04

OFCCP EO 11246 rules were rescinded by final rule on 21 August 2026, so no audit duty is assumed.

05

An EEO-1 rescission reached OIRA on 14 May 2026, so no form is assumed to still be there.

Decide
06

Art. 9(6) puts accuracy on the employer's management, so the agent puts its name to nothing.

07

A gap is explained by a control that sits downstream of it, and the model is held, not shown.

Out
08

A number resists sourcing, and the failure to source it is what goes into the record.

09

Write back to range records only inside the approval boundaries agreed at implementation.

Product statement

Drawing, matching and holding belong to the agent. The number a pay decision rests on belongs to a named reward officer, who approves it and owns where it came from.

Example workflow

One market range, sources to approval

AgentHuman
1Sources receivedPublished survey cuts, job records, internal pay history or a vendor extract
2Matched and datedWhich survey, which cut, how many employers reported into it and the day it was collected
3Range drawn and cells countedThe benchmark, the internal spread, the reporter count and the vintage carried forward
4Controls appliedProvenance checks, cell-count checks, vintage checks and match confidence
No human action required

Stages 1 to 4 run unaided, and no range is approved at any of them — the agent is drawing, and the officer lane opens at the approval gate.

5DecisionSplits at the approval gate
Provenance traced end to end

Goes to the named reward officer to approve.

Anything untraced

Adds a senior reward read first.

Reward review

The range is held with the surveys under it, the reporter counts and the sources it could not trace.

Approve · Attach source · Send to reward review
Approved — by the named reward officer
6Range and job-architecture records updatedOnly where write access and retention policy permit it
7Outcome evaluatedProvenance capture, source currency, officer corrections and what review found
Corrections

Each reward correction is counted in the evaluation.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Approving a range for use in a pay decision.
Confirming the accuracy of a pay-gap report.
Deciding a source is lawful to bring into the corpus.
Setting the market position a job family is paid at.
Automation boundaryAgent acts unaided
Write the survey and its vintage beside the range.
Characterise each source by who collected it and how aggregated it is.
Print how many employers reported into a cut.
Carry the vintage of each survey forward into the range recommendation.
No range is approved except by a named reward officer, inside the agreed boundaries.
Judging whether a job matches a survey benchmark.
Telling a works council a category gap is justified.
Choosing which market a job family is benchmarked to.
Changes to a range, a grade or a pay decision.

Example output

One market range, annotated

This serves a reward team who may have to reconstruct, six months after a pay-gap report was submitted under Art. 10(1)(c), which survey a benchmark came from; below is one range exactly as the agent leaves it.

Market range · single job familyIllustrative example
Range
Recorded as
Survey
Evidence of record
Confidence
Held for
Range review, senior engineering family
Traced to a named survey cut
Reporter count, printed
Survey cut, 3 August 2026
Held unapproved
The named reward officer
As receivedDrawn from a third-party survey cut and the internal job architecture, and it claims nothing past them.
What the record holds Third-party survey Job-architecture record Internal pay ledger
Why no approval hereApproving a range is a judgement that the named reward officer makes.
ActionApproveAttach sourceSend to reward review
What the score decidesBelow the configured threshold a range gets a reward read before the officer sees it.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Each rangeFrom the survey that carries it
03Evidence

Where the evidence is used

The agent does not vouch for a survey, only for who ran it, when it was collected and how many employers stood behind the cut; the same Guidelines reach a tool generating wage recommendations.

01Approved path

Whose numbers are these

Read in the January 2025 Guidelines: an exchange may be unlawful whether or not that effect was intended, and a competitor stands behind the cut.

02Human review

What was checked, and not found

No safety zone for information exchange appears in the January 2025 Guidelines. The 1996 statements that carried one were withdrawn by DOJ on 3 February 2023 and rescinded by the FTC on 14 July 2023, and the 2016 guidance they sat under is still downloadable — availability is not currency.

04Build an evidence trail

The range, the survey it came from and the officer who approved it stay together.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Third-party survey dataSurvey vendors · published cuts
Benchmark ranges by job family
Job architectureGrades · families · job records
Matches and scope of work
Internal pay ledgerPayroll · pay history · headcount
Current pay by job and grade class

Agent

Compensation benchmarking

Reads the sources
Draws the range
Holds for the officer

Reporting and filing recordsPay-gap reports · filing archive
Category cuts and past submissions
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six rules between the survey and the range

Six rules against one edge, the last the finest. What measures true is set out in the map below.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeCut the agent back to reporting the spread when evaluation or production signals degrade.Roll back
L5Vintage monitoringTrack model, prompt and sourcing rules, and note the vintage each range was drawn under.Track
L4Source trailRecord each range, the surveys under it and each read of the file, since an ingestion log outlives the four-year antitrust clock.Record
L3Officer approvalHold the range for a named officer; Art. 9(6) puts accuracy on management, not on a method.Gate
L2Provenance guardrailsTest each category gap against Art. 10(1)(a), and refuse a range drawn under a withdrawn safe harbour.Restrict
L1Aggregation limitsRoute a thin cell or a wide category gap to a senior reward read before the range moves on.Require review
Model coreRange drawn — the surveys, the matches, the reporter counts and the vintages
L1 – L2Test whether a range may stand
L3Leaves the approval to a named reward officer
L4 – L5Keep the range and the survey behind it
L6Reports the spread and no recommendation when signals degrade

How Nestack evaluates it

Evaluate the whole derivation — not only the range that comes out of it.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the range a pay decision uses
Depth of coverage ▼
E1Final-output evaluationDid the range record the surveys it was actually drawn from?
E2Step-level evaluationDid the agent read the right survey, the right cut and the live match?
E3Tool evaluationDid it read and write the correct job family and the correct market?
E4Confidence calibrationDo low-confidence ranges actually attract more officer corrections?
E5Slice evaluationHow does performance change across specific job families?
E6Business outcomeHow many ranges needed a correction before the officer approved?
Floor — the sources a range rests on

Failure modes

Where each failure originates in the agent

Seven failure modes, each pinned to the stage where it first surfaces.

Agent lifecycleDirection of processing →
01 · Retrieval1 mode
UP-03

Tainted source ingested

A shared folder enters the corpus as market evidence.

Stage gathersThe surveys, the vintages and the job matches
02 · Reasoning2 modes
UP-04

Job matched on title alone

A whole survey distribution is imported wholesale.

UP-06

Withdrawn safe harbour read as live

A rescinded survey condition is worked as shelter.

Stage proposesThe survey, the match and the reporter count
03 · Tool / write2 modes
UP-02

Thin cell passed forward

A range moves on without the senior reward read.

UP-05

Range bound to wrong family

The number is filed against another job family.

Stage writesOnly where write access and approval policy allow it
04 · Output1 mode
UP-01

Approved, provenance unrecorded

The record shows approval but no survey beneath it.

Stage returnsThe range a pay paper carries and an offer uses
05 · Change / Version1 mode
UP-07

Silent vintage drift

Aged survey data is re-dated as though newly collected.

Stage tracksModel, prompt, sourcing rules and survey dates
Sev-1 · a range approved on no source Sev-2 · a competitor figure reaches a range Sev-3 · source degrades, range held back

Affected slices

Hot-market families absorb the corrections

A function-level correction figure can read clean while hot-market technical families carry most of the corrections. Nestack reports the correction rate by job family, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Hot-market technical families9.4%3.7× Review
Newly created hybrid roles6.6%2.6× Review
Multi-country job families4.1%1.6× Watch
Established core job families1.9%0.7× Normal
Bar: correction-rate lift vs. core-job-family baseline · scale 0–4.0× · tick marks the 2.0× review threshold 2 of 4 slices over threshold

Evidence-linked improvement

What an untraced number costs

A cycle ends when the number traced to a competitor is a standing case. That suite is what the next review round is measured against.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Correction rate rises on hot-market technical families.

02Diagnose

The benchmark that came from a peer spreadsheet rather than a survey, and read exactly like one, is worked backwards until a single cause is left standing.

03Improve

The range goes out numbered, and the sources beneath it ride with it.

04Verify

One provenance case still failing is enough to hold the round back.

05Learn

The case stays on, and the sourcing rules are redrawn beside it.

Learn → DetectThe return edge. The next range is measured against a suite one case longer.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources, range drawing, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Job-matching review and automation-boundary scoping.
02Survey, job and pay-ledger sources.
03Source-to-range and provenance-coverage mapping work.
04Survey source ingestion.
05Survey, match and job-family binding.
06Cell-count scoring and review routing.
07Reward officer approval workflow.
08Reward-system integration.
09Provenance and matching cases.
10Guardrails and recommendation controls.
11Source-trail instrumentation.
12Deployment, documentation and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne job family, one round ProductionProduction reward workflow AdvancedMultiple families / markets
Introduced at Pilot
Range drawing to your job architecture
Named officer approval
Job-architecture baseline
Introduced at Production
Reporting by job-family class
Reward review workflow in your systems
Approved write-back
Survey-and-ledger integration
Introduced at Advanced
Multi-source provenance rules
Cross-market range packs
Large survey libraries
Multi-market range controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on integrations, workflow complexity, survey volume, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Your live job families and the survey each is benchmarked to Source capture and vintage versioningWeek 1
02Representative survey cuts, job records and pay history Source binding, match logic and the job-architecture baselineWeek 2
03Your pay calendar and the reward officers it names Source mapping, provenance capture and the automation boundaryWeek 1
04Access to relevant APIs, feeds or exports Survey, job and pay-ledger assessment, then integration setupWeek 2
05Ranges you would not want subpoenaed Provenance cases and the market roundWeek 4
06What no range may be called Cell-count scoring, review routing, guardrails and approval controlsWeek 3
07A named reward officer who approves the range Handover to the named reward officer, then pilot and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Widths follow the work rather than the grid, which is why one band is drawn beneath another in week five.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Source discovery, vintage versioning and the automation boundary W2Survey integration and the job-architecture baseline W3Range drawing, match logic and approval controls W4Evaluation suite, provenance cases and the failure round W5Reward-system integration, pilot ranges and targeted corrections W6One pay round run under the reward lead, then Agent Care handover
Reading the bandA bar spans only the weeks its own work is named for, and week five carries two.
At the end of W6When the source record validates, Agent Care takes the agent on.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Reward AI agent

Build a compensation benchmarking agent around the provenance your last range never carried.

Show us one job family you benchmark each year and the survey the last range came from. In Dorrell, No. 1:25-cv-02251 (D. Md.), one defendant settled in May 2026 and the case was then dismissed largely without prejudice on 5 August 2026. Cost arrived before vindication did.

Nestack Agents · Reward benchmarkingAGT-HR-10 · Agent Care available after launch