Nestack Agent Care
Industries / Media & Entertainment / Metadata & archive search

Media AI agent · Archive

Metadata-Tagging & Archive-Search AI Agent

Watch, listen to and read material already in the library, propose timecoded metadata and a search index — with face and voice matches left as candidates, restrictions surfaced, and an archivist accepting the record.

4–6 weeksTypical delivery
Your stackDeployment
Archivist onlyCatalogue record
Agent CareAfter launch

What this agent does

Describes and indexes what the library holds

In
01

Take material already in the library — programmes, rushes, stills and audio — with whatever record is filed against it.

02

Read the existing catalogue entry, the shot list, the paperwork on file and any restriction already recorded.

Reason
03

Watch and listen through the item, mark the shot changes and describe what is in each segment against timecode.

04

Transcribe the spoken word for search, and propose subjects, locations, music heard and identification candidates.

05

Work to the client's thesaurus and description standard, keeping the original catalogue wording beside a re-wording.

Decide
06

Hold an identification the evidence does not carry, and say what the agent was uncertain between.

07

Route restricted items, third-party material and historical wording that needs handling to a named archivist.

Out
08

Return proposed metadata beside the existing record — timecoded, sourced and marked as proposed, not as catalogue.

09

Retain what was read, what was proposed, what the archivist accepted or rejected, and the searches that were run.

Product statement

The agent describes, proposes and indexes. What enters the catalogue of record, and whether a person on screen is named, stays with a media librarian or archivist.

Example workflow

One archive item, end to end

AgentHuman
1Item taken upA backlog reel, an unlogged rushes bin, a new acquisition or a re-request against a record already held
2Existing record readThe catalogue entry, the shot list, the paperwork filed against it, the restrictions recorded and the thesaurus in force
3Content describedShot changes, scene description, spoken-word transcript, subjects, locations and music heard, each against a timecode
4Identification proposedFaces and voices scored against the client's named reference sets as candidates, carrying what each was matched against
No human action required

Stages 1 to 4 run without a person in the loop — the item is described and indexed before anyone is asked to read anything. A restriction on the item ends that stretch.

5DecisionSplits on the evidence behind a name and what the record restricts
Evidence clear, nothing restricted

Reaches the archivist ready to accept.

Thin match or a restriction on file

Goes back named as open, not answered.

Media librarian or archivist

Reads the proposed description beside the item and the record already held, and decides what enters the catalogue.

Accept · Re-word · Send back
Accepted — handed back
6Accepted into the catalogueWritten to the record only where a media librarian accepted it; existing fields are added beside, not replaced
7Outcome reviewedWhat was accepted, re-worded or rejected, what a known-item search failed to return, and corrections by collection
Re-wordings

Every description an archivist rewrites is counted.

What should not run autonomously

Human approval stays in control

Outside the boundary — human approval required8 items
Accepting a face or voice match as identification.
Overwriting or amending an existing catalogue record.
Switching face or voice matching on for a collection.
Clearing restricted or third-party material for re-use.
Automation boundaryAgent acts unaided
Describe shots, scenes, subjects and locations against timecode.
Transcribe the spoken word and index it for search.
Propose face, voice and music matches as ranked candidates.
Surface the restrictions already recorded against an item.
Write actions run only inside the approval boundaries agreed during implementation. The catalogue is not one of them.
Removing or re-wording a restriction already recorded.
Re-writing historical description without a contextual note.
Declaring a search of the library complete.
Changing the thesaurus, label set or match thresholds.

Example output

One archive item, annotated

Everything the agent proposes is attached to the item and the timecode it came from.

Metadata output · single archive itemIllustrative example
Item
Segment
On file
Handed to
Confidence
Named person
News rushes, 1987
10:04:12 – 10:06:48
Shot list only
Archivist — held
88%
Two candidates, neither taken
As receivedThe item, the timecoded segment and whatever the catalogue already holds — nothing here is inferred.
Evidence used Reference set, 40 named Transcript, two mentions Shot list, 1987 entry
Why no name was takenTwo people in the reference set score alike, and the transcript names neither.
ActionAcceptRe-wordSend back
What the score decidesConfidence decides how hard the archivist re-checks the description, not who the person is.

Value

Where AI adds value

The same four claims, placed at the point in the workflow where each one applies.

Where the value landsValue 01 – 04
The library as it standsCatalogued, part-catalogued and unlogged
03Description & indexing

Describe in your own vocabulary

Use the client's thesaurus, description standard, reference sets and the record already filed against the item.

01Approved path

Reach the material nobody logged

Unlogged holdings arrive with a timecoded description and a transcript, so a search can find them instead of depending on who remembers them.

02Human review

Put the doubtful name to a person

Thin matches, restricted items and historical wording reach an archivist as questions rather than entering the catalogue as facts.

04Build an evidence trail

Retain the item read, the proposed description, the reference set behind a match, the archivist's decision and the searches run — on both paths.

Integrations

Typical integrations

Five system groups connect to the same agent. Which of them are in scope is decided in discovery.

Archive & MAMAvid MediaCentral · Dalet · Iconik
Archive catalogues · legacy databases
Storage & proxiesLTO and deep archive · Object storage
Proxy and transcode services
Vocabularies & identifiersIn-house thesaurus · EBUCore · PBCore
Dublin Core · EIDR references

Agent

Metadata tagging & archive search

Describes the item
Indexes for search
Holds the uncertain

Search & researchSearch index · newsroom systems
Production and research tools
Observability & evaluationOpenTelemetry · Langfuse
Supported monitoring/evaluation sources

Integration availability depends on the client's existing systems and API access.

Agent controls

Six layers between the model and your catalogue

Each control wraps the one inside it. A proposed tag clears every layer before an archivist is asked to accept it, and the catalogue of record sits outside all six.

L6 · Outermost — last line of defenceInward → L1 · closest to the model
L6Rollback / safe modeReturn description to your cataloguing team if evaluations or production signals degrade.Roll back
L5TraceabilityRecord the item read, the proposal, the reference set, the decision and the search.Record
L4Scope and restrictionTags are drawn from your thesaurus, and restrictions on file travel with the search result.Restrict
L3Catalogue gateWhat enters the record, and any name attached to a face or a voice, is decided by an archivist.Gate
L2Identity as candidateFace and voice matches are ranked candidates carrying the reference set they were scored against.Rank
L1Evidence bindingEach proposed tag is checked back to the item, the timecode and the source it was read from.Bind
Model coreDescription proposed — shots, subjects, transcript, music heard, match candidates and confidence, all timecoded
L1 – L2Tie every tag to evidence and rank it
L3Decides what may enter the catalogue
L4 – L5Bound the vocabulary, keep the trail intact
L6Pulls automation back when signals degrade

How Nestack evaluates it

Evaluate what was described, what was found and what was missed.

Coverage runs the whole depth of the workflow, and every layer is cut by slice.

Surface — the description an archivist opens
Depth of coverage ▼
E1Final-output evaluationIs the description accurate to what is on screen and in the audio?
E2Step-level evaluationDid it segment the item, transcribe it and map terms to your thesaurus?
E3Tool evaluationDid it read the right item, version and timecode, and write nowhere else?
E4Match and recallIs a proposed name the right person, and what does a known-item search miss?
E5Slice evaluationHow does description change across collections, eras and voices?
E6Business outcomeHow much of the description did an archivist have to re-word?
Floor — the material a researcher can actually find

Failure modes

Where each failure originates in the agent

Seven failure modes plotted against the five stages of the agent lifecycle. The ones that matter most are the ones that name a real person.

Agent lifecycleDirection of processing →
01 · Item & record2 modes
MT-01

Absent restriction reads as none

Nothing recorded is offered as nothing to know.

MT-02

Legacy wording carried forward

A slur in a 1970s entry becomes a live search tag.

Stage gathersThe item, its timecode and the record already on file
02 · Description2 modes
MT-03

Transcript invents a sentence

Speech nobody said is indexed and later found.

MT-04

Regional speech thins the index

The words a researcher would search on are dropped.

Stage describesShots, subjects, the transcript and the music heard
03 · Identification1 mode
MT-05

Confident wrong face match

Two people in the reference set score alike.

Stage matchesFaces and voices against your named reference sets
04 · Index / handover1 mode
MT-06

Empty result read as absence

A search returns nothing and is taken as proof.

Stage returnsThe search index and the description a person reads
05 · Change / Version1 mode
MT-07

Silent thesaurus regression

A vocabulary change re-labels a collection unnoticed.

Stage tracksModel, prompt, thesaurus and reference-set changes
Sev-1 · a tag wrongs a person or a group Sev-2 · a restriction or a holding is missed Sev-3 · description degrades, more is re-worded

Affected slices

A tag is least reliable where the record is thinnest

Early film and tape, rushes with nothing written against them, and the voices least represented in training data are where an archivist re-words most and where a wrong name is likeliest. Nestack reports performance by slice, not only in total.

Slice performance — reported separately, not only in aggregateIllustrative example
SliceFailure rateLift Lift vs. thresholdStatus
Pre-1990 film and tape4.6%2.4× Review
Unlogged rushes, no shot list4.0%2.1× Review
Regional accents and dialect3.3%1.7× Watch
Recently catalogued programmes1.4%0.7× Normal
Bar: archivist re-wording rate lift vs. recently catalogued programmes · scale 0–4.0× · tick marks the 2.0× threshold 2 of 4 slices over threshold

Evidence-linked improvement

A wrong tag outlives the person who wrote it

Metadata is copied outward — into search, into delivery, into other systems. Correcting one catalogue entry does not call back the copies already made.

Improvement cycle · five stagesSwitchback — the path turns at Improve and returns at Learn
01Detect

Re-wordings, rejected matches or a failed known-item search cluster in one collection.

02Diagnose

Traced to the item read, the transcript, the thesaurus mapping, the reference set or the match threshold.

03Improve

The thesaurus entry, reference set or threshold is corrected with your archivist, and version-linked.

04Verify

Re-run against held-out items from that collection, including the descriptions that were rejected.

05Learn

The corrected entry is written back into the vocabulary both your team and the agent describe from.

Learn → DetectThe return edge. Reference sets are re-read each cycle — a name accepted last year is re-checked against the evidence that supported it, not assumed.

Typical build scope

Twelve workstreams across six weeks

The build scope read against the delivery timeline. Week structure follows the six-week plan — discovery, sources and vocabulary, description and indexing, evaluation, integration, then production validation and handover.

Workstream Week 1Week 2Week 3Week 4Week 5Week 6
01Workflow discovery and automation-boundary definition.
02Archive, MAM and proxy-source assessment.
03Thesaurus, label set and description standard.
04Item selection, segmentation and timecode mapping.
05Speech transcription and search indexing.
06Shot, subject and location description.
07Music-heard and audio-cue candidates.
08Restriction and third-party flag carry-through.
09Face and voice reference sets, where lawful.
10Evaluation suite, slices and retrieval tests.
11Archivist review and acceptance workflow.
12Observability, deployment and Agent Care handover.
12 workstreams · 6 weeks · bar shows the weeks a workstream is active — several run in parallel Final scope and sequence confirmed in discovery

Engagement tiers

What each tier includes

Rows are the capabilities named in each tier's scope. Higher tiers include everything below them.

Capability✓ in scope · — not at this tier PilotOne collection, one source ProductionProduction archive integration AdvancedMulti-collection / multi-site
Introduced at Pilot
Timecoded description and segmentation
Speech transcription and search index
Face and voice matches as candidates
Restrictions on file carried into results
An archivist accepts the record
Baseline evaluation
Thesaurus and description standard
Introduced at Production
Archive, MAM and proxy integration
Review workflow and observability
Additional collections and reference sets
Introduced at Advanced
Cross-site, multi-language and enterprise controls
Build price From $5,000 From $8,000 Custom quote
Final build priceConfirmed after discovery based on library size and formats, the state of the existing catalogue, archive and MAM integrations, language coverage, approval controls and deployment requirements.
Separate from buildBuild pricing is separate from recurring Agent Care, which covers managed monitoring, evaluations, incidents and verified improvements after launch.

What we need from you

What you bring, and what we build with it

Each input maps to a piece of build scope and a week in the delivery timeline.

You bringWe build with it
01Access to your archive, MAM and proxy copies Archive, MAM and proxy-source assessmentWeek 1
02Your thesaurus, description standard and house wording Thesaurus, label set and description standardWeek 1
03A sample of items with cataloguing you already trust Item selection, segmentation and the description baselineWeek 2
04How restricted and third-party material is recorded today Restriction and third-party flag carry-throughWeek 3
05Reference sets you may lawfully match against, and who approved them Face and voice reference sets, where lawfulWeek 4
06Searches that came back empty, and the items you know you hold Evaluation suite, slices and retrieval testsWeek 4
07Named archivists and librarians who accept a record Archivist review workflow, then cataloguing under sign-offWeeks 5–6
Nothing else is required Deployment, documentation and Agent Care handover are ours.

Delivery timeline

Four phases across six weeks

Phases are drawn over the weeks they actually occupy. Week 5 carries both the retrieval tests and the first records an archivist accepts.

Phase W1W2W3W4W5W6
Discovery W1
Build W2 – W3
Evaluate W4 – W5
Pilot & Launch W5 – W6
Week focus W1Collections in scope, the description standard and who accepts a record W2Archive, proxy and catalogue sources wired in, and the first items described W3Transcription, indexing and restriction flags carried into the search result W4Reference sets where lawful, evaluation suite and known-item retrieval tests W5Archive integration and the first records accepted under sign-off W6Cataloguing running under your archivists, then Agent Care starts
Reading the bandRestrictions reach the search result in week 3, before any face or voice matching is switched on in week 4. The bars show that order, not a smooth ramp.
At the end of W6Records have entered the catalogue under a named archivist, with the item, the timecode and the evidence behind every tag on file, then Agent Care takes over monitoring.
DurationSix-week plan shown · typical delivery 4–6 weeks depending on scope confirmed in discovery.

Next step · Media AI agent

Build an archive-metadata agent around your catalogue and your archivists.

Show us a collection nobody has had time to describe, the standard you catalogue to, and a search your team ran last week that came back empty. We'll describe part of that collection against your own thesaurus and show you what it proposed, what it refused to name, and what the search still cannot reach.

Nestack Agents · Metadata tagging & archive searchAGT-ME-08 · Agent Care available after launch