Workflows

Searching Broadcast Archives: Making Decades of News Footage Findable

Broadcast stations sit on decades of archive footage that is barely catalogued. Retroactive AI indexing makes old footage valuable again without requiring a migration or re-tagging project.

FrameQuery Team10 April 20265 min read

Every broadcast station has an archive problem. Decades of footage sitting on SANs, nearline storage, and LTO tape libraries. Tens of thousands of hours spanning hard news, features, investigative pieces, weather coverage, live events, and B-roll. The footage is preserved. Finding a specific clip from 2014 requires either exceptional institutional memory or a lucky search through incomplete metadata.

The archive is one of the station's most valuable assets and one of its least used. The footage has already been paid for, and it retains value for future stories, retrospectives, and context pieces if someone can find it.

Most of it cannot be found, because it was never adequately catalogued.

The cataloguing gap

Broadcast archive metadata was typically entered at the point of ingest. A tape operator or editor would log the date, story slug, reporter name, and a brief description. On a good day, the description included key subjects, locations, and topics. On a busy news day, it might include nothing beyond the slug and date.

This inconsistency compounds over years. An archive built by a dozen different operators across two decades has widely varying metadata quality. Some clips have detailed descriptions. Others have a two-word slug. The footage from the year the station transitioned from tape to file-based might have gaps where the workflow was still being figured out.

Searching such an archive by metadata produces unpredictable results, and there is no way to know what a search missed.

The cost of not searching

When archive footage cannot be found, the alternative is to re-shoot or go without. Both have costs.

Re-shooting means sending a crew to capture something that already exists in the archive. An establishing shot of a courthouse. B-roll of a busy intersection. An exterior of a business that has since closed. For a daily news operation, dispatching a crew costs time and resources that could be allocated to the current story. Historical footage cannot be re-shot.

Going without means the story airs with less context. A report on a policy change lacks footage from when the policy was first debated. A profile piece on a public figure omits earlier appearances that would add depth. The story is weaker because the footage was inaccessible.

Stations that can search their archives produce coverage with more context.

Retroactive AI indexing

The traditional approach to making an archive searchable is a cataloguing project: hire archivists or assign staff to review every clip and add detailed metadata. For an archive of 20,000 hours, this is a multi-year effort costing hundreds of thousands of dollars. Most stations cannot justify the expense, so the archive remains under-catalogued indefinitely.

With AI indexing, models analyze the content and generate searchable metadata. Coverage is consistent regardless of archive size.

A retroactive indexing pass applies four layers of analysis to every clip:

Transcription. Every word spoken in every clip is converted to searchable text with timestamps. Interviews, press conferences, reporter standups, live shots, and anchor reads all become searchable by what was said.

Speaker diarization. Each voice is identified and tagged, so searches can be filtered by who said something. In a broadcast archive the same reporters, anchors, and public figures appear across hundreds of clips.

Scene description. AI-generated descriptions of the visual content in each segment. Aerials, establishing shots, close-ups, press conferences, courtroom footage, weather events. The descriptions make B-roll and non-dialogue footage findable.

Face recognition. People are detected and clustered across the entire archive. Once a public figure, reporter, or anchor is identified, every appearance across every clip is linked. This runs on-device, keeping biometric data local.

The output is a search index where every clip is searchable by what was said, who said it, who appeared, and what was shown, from yesterday's footage to clips from 15 years ago.

Practical considerations for large archives

Processing time. FrameQuery processes footage at about five minutes per hour of video:

  • 5,000 hours (small station, five years): approximately 17 days of continuous indexing
  • 20,000 hours (mid-market station, 10-15 years): approximately 69 days
  • 50,000 hours (major market station, decades): approximately 174 days

Indexing runs in the background and can be paused and resumed. Clips are searchable as soon as they are individually indexed.

Format support. Broadcast archives contain a mix of formats depending on the era and the equipment used. Common formats include:

  • MXF (XDCAM, P2, AVC-Intra) - the dominant broadcast acquisition format
  • XAVC/XAVC-S - Sony's newer codec family
  • ProRes (various flavors) - common in Apple-based edit environments
  • DNxHR/DNxHD - Avid edit environments
  • MPEG-2 - older broadcast and DVD-era footage
  • H.264/H.265 MP4 - newer file-based acquisition and screen recordings

FrameQuery decodes all of these natively, with no transcoding step. Transcoding tens of thousands of hours of footage would add weeks or months to the process and a large amount of storage.

Storage access. Archive footage typically lives on SAN storage, a NAS, or connected LTO libraries. FrameQuery reads from any mounted storage location. Point it at the archive volume, and it processes whatever it finds. Files do not need to be moved or copied.

For tape-based archives that have been migrated to file storage, the footage is ready to process as-is. For footage still on LTO, it needs to be restored to disk-accessible storage first. FrameQuery does not read directly from tape.

Index size. The search index is compact relative to the source footage. Expect about 1 to 2 MB of index data per hour of footage. A 20,000-hour archive produces an index of about 20 to 40 GB.

Incremental indexing going forward

Once the backlog is indexed, new footage is handled automatically. Source folder monitoring watches designated storage locations for new files. When today's footage is ingested to the archive, FrameQuery detects it and queues it for indexing. The index stays current without manual intervention.

Every new clip is indexed with transcript, speaker, scene, and face data from the moment it enters the archive.

Making the archive earn its keep

Broadcast archives represent millions of dollars of production investment, plus the ongoing cost of storage. That investment pays off when the content is used in current production.

AI search gives archivists and librarians content-level search that manual cataloguing cannot scale to. For stations without a dedicated archivist, it provides search that would otherwise not exist.

Download FrameQuery to make your broadcast archive searchable.