No metadata, no AI: fixing the foundations first
Ask why a media AI pilot failed and the answer is rarely the model. It's the data underneath: assets with no rights information, five spellings of the same presenter's name, and a taxonomy last updated when tapes were physical. AI amplifies whatever you feed it — including the mess.
The minimum viable foundation
You don't need a five-year data program before touching AI. You need four things on the collection you care about: unique and persistent asset IDs, rights and usage metadata a machine can evaluate, transcripts or text proxies for audiovisual content, and a controlled vocabulary for the entities that matter to your business — people, places, teams, programs.
Let AI fix the data, too
The good news is circular: modern models are excellent at cleaning the very metadata they need. Entity resolution, automatic tagging, transcript generation, and taxonomy mapping can be largely automated, with humans reviewing samples rather than every record. We routinely see enrichment backlogs that looked like years of manual work shrink to months.
Sequence it right
Run foundation work and the first AI use case in parallel on the same narrow collection, not as separate phases. The use case keeps the data work honest — you clean exactly what the application needs, and nothing speculative.
Want an honest read on your data readiness? Get in touch.