Find the cost before you buy anything
Every media archive we have looked at has the same story and the same missing number. The librarian who knew where everything was has retired or is one person for forty thousand hours. The naming convention changed three times and the oldest material was digitised by whoever had a deck. Producers search once, find nothing, and work around it. The cost is real and invisible: footage shot twice, stock licensed that the company already owns, a sponsor clip that could not be cleared in time, a documentary team paying an outside archive for material sitting on a shelf upstairs.

So start with two weeks of counting rather than a vendor demonstration. Ask five producers to list, from memory, the last three times they needed something they believed existed and could not find it, and what they did instead. Pull the last year of external stock and archive licensing invoices and ask how much of it duplicates owned material. If your system logs searches, pull the queries that returned nothing and the queries where the user opened no result, which is the same failure with a friendlier log entry.
That exercise ends some projects, and that is a good outcome. A three-thousand-asset library with one active team and a working folder structure does not need a tagging programme; it needs a naming convention and someone capturing rights at ingest. The work below is aimed at archives where nobody can hold the contents in their head any more, which in practice starts somewhere in the low tens of thousands of assets or when more than about ten people search independently.
You are probably here because
- Someone reshot material the company already owned, and it was noticed this time
- The one person who knew the archive is leaving, and the handover is a spreadsheet
- A tagging vendor returned ten thousand labels reading person, indoor and text
- Legal cannot say whether a clip can be reused, so it is treated as unusable
The fourth is the expensive one. An asset nobody can clear is not in your library in any sense that matters.
What a machine can honestly tag
Automated tagging is four different capabilities with very different reliability, and treating them as one product is how archives end up disappointed. Ranked by value returned per dollar spent, on a typical mixed-content library:
| Capability | What you get | How reliable | Honest use |
|---|---|---|---|
| Speech transcription | Timecoded words, speaker turns | Very good on clean speech; degrades on crowd noise, accents, overlap | The backbone of search. Do this first |
| On-screen text | Lower thirds, slates, signage, scoreboards | Good on graphics, poor on motion and low resolution | Names, dates, locations you would otherwise never recover |
| Shot and scene detection | Cut points, keyframes, duplicate segments | Reliable; a solved problem | Navigation, dedupe, finding the reused b-roll |
| General visual labels | Objects, settings, activities | Broadly correct and rarely useful | Weak filters at best. Rarely worth the storage |
| Known-face matching | Named people from a curated reference set | Good within a defined roster; poor open-set | High value, and it needs a policy before a pipeline |
| Description by a vision model | A sentence per keyframe | Fluent, occasionally confidently wrong | Searchable text, always attributed as generated |
Two of those rows deserve elaboration, because they are where money is usually misspent.
Speech is the most valuable thing in the building
For any library containing interviews, commentary, presentations, sermons, lectures, panels or press conferences, timecoded transcription is worth more than every visual capability combined. It converts an opaque hour of video into a searchable document, it lets someone jump to the second where a phrase was said, and it produces the quotes and clip candidates that people are actually hunting for. Transcription cost has fallen far enough that processing a large back catalogue is now a budget line rather than a capital project, and it is the one place where reprocessing the entire archive is usually justified on its own.
Do it properly. Keep word-level timecodes rather than paragraph blocks, or the transcript is a document instead of an index. Keep speaker separation even if you cannot name the speakers, because a producer looking for what the coach said can filter by turn. Store the transcript beside the asset in a form other systems can read, not locked inside a player. And retain the confidence scores, because a search hit in a passage the recogniser was unsure about deserves to be shown differently from one in clean studio audio.
Where it fails is predictable and worth stating in advance: overlapping speech, heavy crowd noise, poor lavalier placement, strong regional accents, and specialised vocabulary. Names are the worst case and also the highest-value target, so supply a custom vocabulary of your people, places, products and recurring terms. That single step often does more for retrieval quality than any model upgrade.
Generic visual labels are usually worse than nothing
Off-the-shelf visual tagging returns thousands of labels like person, outdoor, sky, building, text. They are accurate. They are also useless, because nearly every asset gets them and a tag that matches nine-tenths of the library is a tag that filters nothing. Worse, they crowd the interface, they make the tag list untrustworthy, and they train users to ignore tags entirely — which is a hard habit to reverse when you later add tags that would have helped.
The exception is a narrow, domain-specific label set that you define and validate: your product line, your uniforms, your venues, a small number of shot types like aerial, interview, crowd, empty room. Fifty labels chosen by the people who search, measured against a hand-checked sample, beat five thousand generic ones. Building that set is a taxonomy exercise involving your librarians, not a procurement exercise involving a vendor.
Retrieval value per dollar — how we rank the options for a mixed archive
Our ranking for a mixed library with substantial spoken content. A stills-only or wordless-footage archive reorders this considerably.
The most valuable tag is whether you may use it
Descriptive metadata tells you what is in an asset. Rights metadata tells you whether it can leave the building, and it is the field that decides whether the asset has any value at all. A perfectly tagged clip whose talent release cannot be located is not an asset; it is a liability with good search terms.
The fields worth capturing per asset, and per segment where a segment differs: who shot it and under what agreement, which people appear and whether a release exists with a scan attached, licensed music and its term and territory, location permissions, third-party material embedded inside your own edit, any embargo or expiry date, and the specific approved uses. Then present clearance as a status on the search result itself, so a producer sees before opening the file whether it is clear for broadcast, clear internally, restricted, or unknown.
Unknown deserves to be its own state. Most archives have a large share of material where nobody has ever checked, and collapsing that into either cleared or restricted is a decision being made by omission. Showing it as unknown, with the fields that are missing, turns an unusable pile into a queue somebody can work down in priority order — starting with the material people keep asking for.
Every generated tag records where it came from
Store the source, the model and version, the date, the confidence, and whether a person has confirmed or corrected it. Without that you cannot reprocess selectively when a better model arrives, you cannot tell a human-verified fact from a machine guess, and a librarian's correction gets silently overwritten by the next batch run. That last failure ends internal trust in the system permanently, and it is very hard to earn back.
Send us fifty real searches that failed.
The queries your team typed, what they hoped to find, and what they did instead, to contact@precisionfederal.com. You get back which of them transcription alone would have answered, which need rights data, which need a controlled vocabulary, and which no tagging system will ever solve. One business day, no charge.
contact@precisionfederal.comControlled vocabulary, free tags, and the middle that works
Two failure modes sit at either end. A strict taxonomy designed by committee gets ignored, because it requires the person ingesting material at eleven at night to make eleven decisions correctly. Free tagging produces forty spellings of one concept and no way to browse. Neither is a technology problem, and neither is fixed by buying something.
What works in practice is a small controlled vocabulary for the handful of fields people filter on, with everything else living in free text that search covers. Filters need discipline: programme or series, genre, event, venue, named people, rights status, technical format. Everything else can be prose. Machine tags land in a separate namespace from human tags so nobody has to guess which is which, and promotion from machine to confirmed happens when a person accepts it.
Give synonyms a real home. When a producer types the old name of a venue, the new name, an abbreviation or a nickname, all of them should find the same material. A synonym list maintained by librarians is a modest piece of engineering that outperforms a great deal of clever retrieval, and it is the sort of thing an in-house team can own for years.
Search is a recall problem, and it needs a real test set
Consumer search is tuned for precision because there are a million acceptable answers. Archive search is the opposite: there may be exactly one usable shot in the building and missing it costs a reshoot. Tune for recall, present results densely so scanning is cheap, and make thumbnails and scrub previews good enough that a producer can dismiss a wrong result in a second.
Hybrid retrieval is the sensible default. Keyword matching handles names, codes and exact phrases, which is a large share of professional queries and where semantic search is weakest. Semantic matching over transcripts and descriptions handles the vaguer queries people actually type, like the interview where she talks about the flood. Run both and merge, rather than choosing.
Then build a test set of about a hundred real queries with known correct answers, gathered from your own team. Measure whether the right asset appears in the first page. Rerun it after every change to the pipeline, the vocabulary or the model. Without it, every future change is a matter of opinion, and search quality drifts in whichever direction the last person to touch it preferred.
Fix the front door before the back catalogue
Retroactive tagging is expensive and endless. Ingest tagging is cheap and permanent. Whatever else you do, make sure material arriving this quarter carries its rights, its production context, its people and its consistent identifiers from the moment it lands, because everything captured at ingest is something nobody ever has to reconstruct.
For the back catalogue, resist the urge to process everything at once. Process what people search for. Order the archive by demand: recent years, the flagship programmes, whatever the failed-query log points at. Run transcription across the highest-demand tier first, measure whether search improved, then buy the next tier. An archive-wide processing run before anything has been validated is how a budget disappears into a library nobody was going to search anyway.
What we would not build
- A full-archive tagging run before a demand-ordered tier has proven the pipeline
- Generic visual labels at scale, which fill the interface and filter nothing
- Face recognition without a written policy on whose faces, retention, and who can query it
- A taxonomy with more than about a dozen mandatory fields at ingest, which will be filled with defaults
- Generated descriptions presented as fact, indistinguishable from a librarian's note
- A pipeline whose output overwrites human corrections on the next run
- Search changes with no test set, which makes every improvement unverifiable
What a first build looks like
Twelve weeks, in the order we would do it
Step two is the step that gets skipped and the one that determines whether any of it holds. If the same programme exists under three identifiers across the asset manager, the storage tier and a production spreadsheet, then tags attach to one of them and searches run against another. Every archive we have opened has had some version of this, and it is always older and more tangled than the team believes.
Before you commission anything
- You have counted the reshoots and the outside licensing that duplicates owned material
- Every asset has one identifier and one authoritative record
- Transcription is timecoded at word level and stored outside the player
- A custom vocabulary of your names, places and products is in place
- Rights status appears on the search result, with unknown as its own state
- Machine tags and human tags live in separate namespaces with provenance
- A hundred-query test set exists and is rerun after every change
- Corrections survive the next processing run, provably
- The back catalogue is processed in demand order, tier by tier
Bottom line
Tagging a media library is mostly not a computer vision project. It is transcription, identity, rights and search behaviour, in that order, with a small hand-built vocabulary where filtering genuinely helps. The visual capabilities that dominate the sales conversation contribute least, and the field that decides whether an asset is worth anything is usually recorded nowhere. Measure the reshoots, fix the front door, process what people actually search for, and keep every machine-generated claim labelled as one.
Frequently asked questions
It depends entirely on which capability. Transcription on clean speech, on-screen text from graphics, and shot detection are all dependable enough to index against directly. General visual labels are usually accurate and unhelpful, because they apply to most of the library. Named-face matching works within a curated roster and poorly outside it. Treat every generated tag as a search aid carrying provenance rather than as a catalogue record.
Process in demand order and buy the next tier only after search has measurably improved. Most archives have a long tail nobody searches, and spending the same amount per hour on it as on last season's flagship material is how the budget disappears. Rank by failed queries, recent production need and licensing history, then run the top tier first and re-measure before continuing.
Often not as a separate system. Several mainstream search engines and databases now handle vector similarity alongside keyword matching well enough for libraries in the tens of millions of segments, and running one system instead of two saves real operational effort. The important design choice is hybrid retrieval, merging exact matching for names and codes with semantic matching for vague queries, rather than which storage engine holds the vectors.
Start with the policy, not the model. Decide whose faces may be enrolled, on what basis, how long references are retained, who may run a query and what is logged. Matching against a curated roster of staff, talent or public figures who appear routinely is both more useful and more defensible than open-set recognition. Requirements differ by jurisdiction and some places regulate biometric data specifically, so this is a question for counsel before it is a question for engineering.
No, and the projects that assume it will tend to fail. The pipeline produces volume; librarians produce the vocabulary, the synonyms, the corrections and the judgement about what a piece of material is actually for. The realistic change is that their attention moves from typing descriptions to curating a controlled vocabulary and clearing the rights backlog, which is both higher-value work and the part no model can do.
