Nestack Agent Care
Industries / Media & Entertainment / Moderation & trust safety

Media AI agent · Trust & safety

Content-Moderation & Trust-Safety AI Agent

Triage reported posts, clips, comments and streams against your published policy, assemble the context a reviewer needs and draft the statement of reasons — enforcement against a person stays with your trust-and-safety team.

4–6 weeksTypical delivery
Your stackDeployment
Reviewer onlyEnforcement
Agent CareAfter launch

What this agent does

Triages the report, a person enforces

In
01

Take the report or the detection — a post, clip, comment or livestream segment — with who raised it and on what ground.

02

Read the policy as published, the version in force, and the account's prior strikes, open appeals and standing.

Reason
03

Match the material to a policy clause in the language it was written in, and record which clause and which version.

04

Gather what a reviewer needs to decide — the thread around it, the caption, the account history and the reporter's claim.

05

Mark what it could not read — an unsupported language, a dialect, spoken audio or satire — as undetermined, not clean.

Decide
06

Route the case to the queue that matches it, and send a child-safety or imminent-harm match to the named path now.

07

Hold a first strike, a suspected minor, a protected category or a low-confidence match for a person to decide.

Out
08

Draft the statement of reasons and the appeal route for the reviewer to check, correct, sign and send.

09

Retain the material, the policy version, the evidence, the reviewer's decision and every appeal outcome.

Product statement

The agent triages, gathers and drafts against the policy you publish. Removing content, restricting an account and answering an appeal are decided and signed by a reviewer, and a breach of your policy is not a finding of law.

Example workflow

One report, end to end

AgentHuman
1Report or detection receivedA user report, a proactive detection or a trusted-flagger notice, with the ground it was raised on
2Policy and account readThe clause list as published, the version in force, and the account's strikes, appeals and standing
3Context assembledThe material itself, the thread around it, the caption, the language detected and the reporter's claim
4Clause match proposedThe clause it appears to breach, the confidence on that match, and what the agent could not read
No human action required

Stages 1 to 4 run without a person in the loop — the case is assembled before anyone is asked to read anything. A child-safety or imminent-harm match ends that stretch on the spot.

5DecisionSplits on the category, the confidence and whether a duty attaches
Routine category, clause matched

Reaches the reviewer queue with the notice drafted.

Duty attaches, or nothing could be read

Goes to the named specialist now, not to a queue.

Trust-and-safety reviewer

The reviewer reads the material, the clause, the account history and the drafted notice, then decides the action and signs the reasons given for it.

Uphold · Leave up · Escalate
Signed — handed back
6Action recorded, notice sentOnly the action a reviewer signed, with the reasons, the evidence kept and the appeal route attached
7Outcome evaluatedReviewer agreement, appeal reversals, escalation timing and accuracy by language and dialect
Appeals

Every decision overturned on appeal is counted against the match that caused it.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Removing content or disabling access to it.
Suspending, restricting or banning an account.
Deciding an appeal against an enforcement action.
Signing the notice a user receives after enforcement.
Automation boundaryAgent acts unaided
Assemble the material, the thread, the account history and the report.
Match against the version in force and name the clause.
Mark what it could not read as undetermined, and say why.
Route the case to the matching queue and draft the notice for a reviewer.
Write actions run only inside the approval boundaries agreed during implementation. Enforcement is not one of them.
Handling child sexual abuse material or imminent harm.
Referring a user to law enforcement or a regulator.
Deciding whether material is illegal in a territory.
Changing a policy clause or an enforcement threshold.

Example output

One reported item, annotated

Everything the agent proposes is attached to the report and the account it came from.

Triage output · single reported itemIllustrative example
Reported item
Language
Policy
Proposed route
Confidence
Notice
Short video with caption
Regional dialect, spoken
Version in force
Senior reviewer — nothing actioned
58%
Drafted, unsent
As receivedThe item, the language it was spoken in and the policy version it was read against.
Evidence used Clause and version, dated Thread and caption Account's prior strikes
Why nothing is actionedThe dialect is one the classifier is weakest on, and the account has no prior strike.
ActionUpholdLeave upEscalate
What the score decidesIt decides who reads the case and how soon — never whether the material is illegal.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
Every report and detectionFrom users, detection and trusted flaggers
03Policy match

Apply the policy as published

Read the client's own clause list, the version in force and the account's standing under it, and record which clause was matched.

01Approved path

Hand the reviewer a finished case

The material, the clause, the thread and the account history arrive together — as much as the decision needs and no more — so a reviewer's minutes go on deciding rather than on assembling.

02Human review

Move the reportable case first

A child-safety or imminent-harm match leaves for the named escalation path immediately, rather than waiting its turn behind routine reports.

04Build an evidence trail

Retain the material, the policy version and clause, the evidence, the confidence, the reviewer's decision, the notice sent and the appeal outcome — on both paths.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Content & community surfacesUGC platform and CMS · comments and chat
Livestream and VOD · direct messaging
Detection & matchingHash-matching services · classifier APIs
Language identification · transcription
Case & queue managementReview queue tooling · case management
Ticketing · appeal workflow

Agent

Content moderation & trust safety

Assembles the case
Matches the clause
Routes and drafts

Escalation & disclosureNamed escalation path · preservation store
Transparency reporting · regulator notice
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and the enforcement

Each control wraps the one inside it. A case clears every layer before a reviewer opens it, and the enforcement itself sits outside all six.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeReturn triage to your moderation team if evaluations or reversals degrade.Roll back
L5TraceabilityRecord the policy version, the evidence, the decision and the appeal.Record
L4Reasons and appealThe clause, the facts and the route to appeal are drafted with the case.Attach
L3Reviewer gateAction against an account and the answer to an appeal are a person's.Gate
L2Undetermined handlingWhat it could not read is undetermined, not clean; a reviewer decides.Mark
L1Escalation routingChild-safety and imminent-harm matches take the named path, not a queue.Escalate
Model coreCase proposed — the clause matched, the evidence gathered, the confidence, and the queue it suggests
L1 – L2Decide what a match may claim
L3Decides who acts against a person
L4 – L5Keep the reasons and the trail attached
L6Pulls automation back when signals degrade

How Nestack evaluates it

Evaluate the decision and the reasons given for it.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the case a reviewer opens
Depth of coverage ▼
E1Final-output evaluationDid the clause matched agree with a reviewer reading the same item?
E2Step-level evaluationDid it read the version in force and the account's actual history?
E3Tool evaluationDid it reach the right queue and preserve what had to be preserved?
E4Appeal and reversalWhat came back overturned, on what ground, and how long did it take?
E5Slice evaluationHow does accuracy change by language, dialect and cultural context?
E6Business outcomeHow long did a reportable case wait, and how much did a reviewer rewrite?
Floor — the enforcement a person lives with

Failure modes

Where each failure originates in the agent

Seven failure modes plotted against the five stages of the agent lifecycle. The ones that cost most act against a person.

Agent lifecycleDirection of processing →
01 · Report & context2 modes
CM-01

Coordinated reports read as signal

Mass reporting pushes a lawful account up the queue.

CM-02

The quote read as the abuse

The person reporting harassment is actioned for it.

Stage gathersThe item, the thread, the account and the report
02 · Clause match2 modes
CM-03

Dialect matched to a slur list

A regional variant reads as the banned term.

CM-04

Atrocity evidence read as gore

Human-rights documentation scores as graphic violence.

Stage matchesThe clause, the version in force and the confidence
03 · Routing1 mode
CM-05

Reportable case left in a queue

A child-safety match waits behind routine reports.

Stage routesThe queue, and the path a reportable case takes
04 · Notice & appeal1 mode
CM-06

A clause number and no facts

The user cannot tell what they did, or appeal it.

Stage returnsThe statement of reasons and the route to appeal
05 · Change / Version1 mode
CM-07

Threshold moves, old strikes stand

Yesterday's bans sit under a line nobody re-ran.

Stage tracksModel, prompt, policy-version and threshold changes
Sev-1 · a duty the platform owes is not met Sev-2 · the wrong person is enforced against Sev-3 · review load rises and decisions slow

Affected slices

Wrong enforcement is concentrated, not spread

A policy applied consistently overall can still be applied worst to the users with the least recourse — the languages and formats carrying most of what appeals overturn. Nestack reports performance by slice, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Low-resource languages8.0%5.0× Review
Dialect and code-switched posts4.3%2.7× Review
Livestream and spoken audio2.9%1.8× Watch
Routine reported spam1.1%0.7× Normal
Bar: appeal-reversal rate lift vs. routine reported spam · scale 0–4.0× · tick marks the 2.0× threshold 2 of 4 slices over threshold

Evidence-linked improvement

A reversal arrives after the person has served it

An appeal upheld six weeks later does not give back the six weeks. The cycle is judged on how fast the pattern behind the case is found.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Appeal reversals, escalation delay or reviewer disagreement move in one cohort.

02Diagnose

Traced to the context assembled, the clause matched, a term list, the threshold or the routing rule.

03Improve

The clause mapping, term list or routing rule is changed with your trust-and-safety lead, and version-linked.

04Verify

Re-run against held-out reports from that cohort, including the ones an appeal overturned.

05Learn

The overturned case is kept beside the clause it was wrongly matched to.

Learn → DetectThe return edge. A change to a clause changes what you have told users you will do, so it is published before it is enforced, not after.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, policy and escalation, triage and routing, evaluation, review tooling, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Workflow discovery and automation-boundary definition.
02Policy, clause and enforcement-ladder map.
03Escalation and preservation, named owners.
04Report, detection and content-source links.
05Context assembly and clause-matching logic.
06Language and dialect coverage assessment.
07Queue routing, priority and hold rules.
08Statement-of-reasons drafting and appeal workflow.
09Evaluation suite, slices and regression cases.
10Reviewer console and exposure controls.
11Transparency and enforcement reporting.
12Observability, deployment and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne surface, one policy area ProductionProduction queue integration AdvancedMulti-surface / multi-language
Introduced at Pilot
Case triage against your published policy
Escalation and preservation routing
Reviewer exposure controls
Enforcement and appeals decided by a person
Statement of reasons and appeal route
Baseline evaluation
Queue routing, priority and reviewer console
Introduced at Production
Detection and hash-matching connections
Observability and slice evaluation
Additional languages and dialect coverage
Introduced at Advanced
Transparency reporting and enterprise controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on the surfaces and languages in scope, report volume, detection and case-management integrations, review controls, reporting depth and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01The policy you publish, and the clause list behind it Policy, clause and enforcement-ladder mapWeek 1
02Your path for child safety and imminent harm, and who owns it Escalation and preservation, with named ownersWeek 1
03Access to your content, report and case-management systems Report, detection and content-source linksWeek 2
04The languages and dialects your users actually post in Language and dialect coverage assessmentWeek 3
05Enforcement decisions you had to reverse on appeal Evaluation suite, slices and regression casesWeek 4
06How your reviewers are protected from what they are shown Reviewer console and exposure controlsWeek 5
07Named trust-and-safety reviewers and an appeals owner Appeal workflow, then supervised cases and production validationWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Phases are drawn over the weeks they actually occupy. Week 5 carries both the language slices and the first cases a reviewer works from the agent's triage.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Policy, clause ladder and the escalation path with its named owner W2Source connections, context assembly and the triage baseline W3Clause matching, language coverage and queue routing W4Reasons drafting, appeal workflow and the evaluation slices W5Reviewer console, exposure controls and supervised cases W6Transparency reporting, production validation and Agent Care handover
Reading the bandThe escalation path is built and walked through in week 1, before the agent triages anything. A child-safety or imminent-harm case does not wait on the evaluation phase.
At the end of W6Reviewers have worked live cases from the agent's triage and signed the reasons on everything that went out, then Agent Care takes over monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Media AI agent

Build a moderation agent around your policy and your reviewers.

Show us the policy you publish, a week of the reports you actually get, and the enforcement decisions you had to reverse. We'll triage that week against your own clauses and show you the cases it refused to decide, the languages it could not read and the notices a reviewer would have had to rewrite.

Nestack Agents · Content moderation & trust safetyAGT-ME-10 · Agent Care available after launch