Media Archive Preservation Software: Which Carriers Do You Digitize First, and Can You Legally Sell Any of Them?
$80,000 to $170,000 for a first release in 14 to 20 weeks covering the item and carrier register, condition assessment, risk based prioritization and vendor batch workflow with chain of custody, based on Digital Heroes delivery experience. A full platform adding ingest with checksum verification and technical metadata extraction, fixity scheduling, storage migration planning, per item rights clearance and links into your media asset management system runs $200,000 to $500,000 across 9 to 15 months. Build when your holdings run past roughly 50,000 physical items across mixed carrier types with a funded multi year program. Do not build if you hold a few thousand born digital files: a preservation repository product handles that properly and cheaper.
Why a digitization program stalls in its second year
Year two of a five year funded program. The vault holds around 180,000 items across 2 inch quad, 1 inch Type C, U matic, Betacam SP, Digital Betacam, DAT, HDCAM, 16mm and 35mm negative and print, plus a corner of nitrate that is stored separately and thought about rarely. The digitization queue is ordered by title recognition, because that is the only ranking anyone could agree on. A vendor batch of 40 tapes comes back with three unplayable, marked on the packing note as binder issues, and nobody recorded which deck was used or whether baking was attempted. The rights researcher then reports that the highest value series in the queue has music cues that were cleared for original transmission only, so the files just made cannot be sold or streamed.
The money was spent. The decay clock did not pause for any of it. Acetate stock keeps hydrolyzing whether or not it is at the top of your list, magnetic binders keep absorbing moisture, and the pool of engineers who can service a working quad machine keeps shrinking. The real constraint in a moving image archive is not storage cost. It is that the intersection of a playable carrier, a working deck and a person who can operate it is closing on a schedule you do not control.
Meanwhile the tooling is split. A collections database that was built for documents holds catalog records. A spreadsheet holds the digitization queue. The vendor holds the chain of custody in their own system. Checksums live in a text file next to the files. Rights live in contracts and a researcher's notes. Nobody can answer the only question that matters to a funder or a board: for a given amount of money and time, which items should we do next, and what will we be legally able to do with them afterward.
Problem 1: condition data does not exist in a form you can prioritize on
Prioritization needs three inputs multiplied together: how fast the carrier is dying, how valuable the content is, and how clearable the rights are. Most archives have a partial answer to the second and nothing structured for the first or third.
Condition is format specific and that is exactly why generic collection systems cannot hold it. Acetate film needs an acidity reading, typically taken with A-D strips, plus shrinkage measurement and splice condition. Magnetic tape needs a binder condition judgment, a note on whether baking was performed and at what temperature, and an assessment of edge damage and pack quality.
A custom build gives each carrier type its own condition schema, captured on a tablet in the vault with barcode scanning, and then computes a decay risk score that feeds prioritization. The output is a queue you can defend in a funding meeting, because it shows why a mediocre regional news series from 1976 outranks a beloved drama from 1991. The news is dying first.
Problem 2: vendors hold the chain of custody and give it back as a spreadsheet
Digitization at scale is outsourced in batches. Items leave the vault, cross a loading dock, sit in a vendor facility for weeks, get transferred, get quality checked, and come back with files on a drive or an LTO tape and a manifest in whatever format the vendor uses. If the archive's record of that journey is an email trail, then a lost item, a mislabeled reel or a failed transfer becomes an archaeology exercise.
What a build does is make the batch a tracked object with a manifest generated by the archive, not by the vendor. Every item has a scan out and a scan in. The vendor gets a portal or an API to post transfer results, QC notes, deck used, operator and any interventions performed, and those attach to the item permanently. When a file is questioned in eight years, the provenance is there: this specific carrier, this condition, this deck, this operator, this date, this checksum. That record is also what makes the resulting file trustworthy as a preservation master rather than just a copy.
Problem 3: preservation repositories start after the file exists
Preservica and Arkivum are legitimate products and we would not talk anyone out of them for what they do. They are digital preservation repositories aligned to the OAIS model, strong on fixity checking, format identification, migration of digital objects and audit trails for records. If your problem is holding born digital material safely for decades with defensible custody, they solve it.
Their model begins where the hardest part of a moving image archive ends. They do not run your vault, your shelf locations or your carrier condition. They do not schedule scarce playback equipment or manage vendor batches and chain of custody across a loading dock. Their metadata models come from the records and documents world, so carrier specific technical detail and per item rights clearance sit awkwardly if they sit at all. And they do not connect an item to the commercial channel that pays for the program, which is the link that keeps a preservation budget funded.
The pragmatic answer is usually not to replace them. It is to build the operational layer in front, covering condition, prioritization, vendor workflow and rights, and let the repository do storage, fixity and format management underneath. A build that ignores that division and reimplements a preservation repository from scratch is wasting your money.
Problem 4: rights clearance decides whether any of this can be monetized
An archive earns its budget by being useful. Useful means licensable, streamable or exhibitable, and that is a per item legal question with several independent answers that must all come back yes. Music cues cleared for original broadcast only. Performer and contributor consents that predate any concept of digital distribution. Third party footage licensed for a term that expired decades ago. Underlying literary rights. Contributor consent for people who did not know the interview would outlive the program.
Most archives keep this in a researcher's head and a filing cabinet, so the same series gets researched three times by three people over ten years. A build turns clearance into a structured status per item with the evidence attached: the cue sheet, the contract page, the consent form, the territory and term, and an expiry that triggers a review. Then the digitization queue can weight clearability, so you stop spending scarce vendor time on material you will never be permitted to show.
There is one place machine assistance genuinely earns its keep here. Cue sheets, contracts and consent forms exist as scanned paper in wildly inconsistent layouts, and document extraction can pull the parties, the term, the territory and the usage type into structured fields for a human to confirm. It does not replace the researcher. It stops the researcher from retyping.
Problem 5: fixity and migration are a schedule, not an event
Once the files exist, the work changes shape but does not stop. Preservation masters need periodic fixity verification against stored checksums so silent corruption is detected while a second copy is still good. Storage tiers need planning, and if you use LTO you are on a hardware clock, because each drive generation reads back only a limited number of previous generations and writes back fewer. That means every tape you write today has a known migration horizon whether you have budgeted for it or not.
A build keeps a storage plan per copy: where each copy lives, on what medium, written when, verified when, and when it must be migrated. The report a board actually needs is one page. How many items exist in preservation master form, how many copies of each, where they are, when they were last verified, and what the migration bill looks like over the next five years. Almost no archive can produce that page today without a week of work.
What this costs and how long it takes
A first release covering the item and carrier register with barcodes and shelf locations, format specific condition capture, risk based prioritization and the vendor batch workflow with chain of custody runs $80,000 to $170,000 and ships in 14 to 20 weeks in our delivery experience. A full platform adding ingest with checksum verification and technical metadata extraction using tools such as MediaInfo and FFprobe, fixity scheduling, storage and migration planning, per item rights clearance with document extraction, and integration with a media asset management system and sales channels runs $200,000 to $500,000 across 9 to 15 months.
What drives price up specifically in archives: the number of distinct carrier types, since each needs its own condition schema and its own vocabulary. Retrospective data cleanup, which is almost always underestimated because the existing catalog has decades of inconsistent title, series and episode identification. Integration with a media asset management platform such as Dalet, Avid or Vidispine, where the hard part is agreeing which system owns which field. And rights document extraction volume, since ten thousand scanned cue sheets is a different project from five hundred.
What keeps it down: starting with one collection, one carrier family and the prioritization model. If the first release only tells you what to do next and proves it, it has already changed how the money is spent.
Build versus buy, and when buying is right
Buy if your holdings are mostly born digital, in the low thousands of objects, and your obligation is safe custody with defensible audit. Preservica or Arkivum will serve you better than anything we would build, and we would say so on the first call. Buy as well if you have no digitization program and no funding for one, because operational software for a program that does not exist is an expensive way to describe an intention.
Build when two or more of these are true. You hold more than roughly 50,000 physical items across mixed carriers with real decay risk. You have multi year funding and have to defend the sequencing of that spend to a board or a public funder. You use external digitization vendors and cannot currently reconstruct chain of custody for a batch from last year. Rights uncertainty is blocking monetization of material you have already digitized, which is the most painful version because the money is already spent. Or you run more than one vault and items move between them.
The honest tipping point is the prioritization question. If you can answer what to digitize next, in what order, and why, with evidence, you probably do not need this. Almost nobody can.
How to choose a developer for archive and preservation software
Ask them to whiteboard the data model before you sign. A developer who understands this domain separates the intellectual work, the manifestation, the physical carrier and the digital instantiation, and knows that one work can have six carriers of differing quality and three digital copies of differing provenance. A developer who draws assets and files has built a document store and will discover the difference on your program.
Ask how they would handle condition data across carrier types. If the answer is a free text notes field, they have not thought about prioritization, because you cannot rank on prose.
Ask what they will integrate rather than rebuild. The right answer usually keeps a preservation repository for storage, fixity and format management, keeps your media asset management system for access copies and editorial metadata, and builds the operational layer that neither one covers. A developer who proposes to replace all of it is quoting a project, not solving a problem.
Ask who owns the code and put it in the contract before kickoff. You should own the repository, the cloud accounts and the right to hire anyone else. At Digital Heroes the client owns the code from the first commit. Archives think in decades, and a system that outlives its original vendor relationship is the only kind worth building here.
The evidence behind this guide
Independent findings on why this investment pays off. Every link goes to the primary source.
- The median annual wage for U.S. software developers was $133,080 in May 2024, and employment is projected to grow 15% from 2024 to 2034 - a core input to any in-house build-vs-buy TCO model. Source: U.S. Bureau of Labor Statistics (2024) →
- Companies in the top quartile of McKinsey's Developer Velocity Index had 2014-18 revenue growth four to five times faster than bottom-quartile peers, showing that software-building capability is a driver of business performance, not just a support function. Source: McKinsey & Company (2020) →
- In a McKinsey global survey of 1,259 respondents, only about 20% said their organizations excel at decision making, and just 37% said their organizations' decisions were both high quality and high in velocity. Source: McKinsey & Company (2019) →
- Qualtrics research (Q3 2023 survey of ~28,400 consumers across 26 countries) estimated bad customer experiences put roughly $3.7 trillion in global revenue at risk annually, a 19% jump from the prior year's $3.1 trillion; 64% of customers say they will switch companies over poor service regardless of how much they like the product. Source: Qualtrics XM Institute (via Forbes) (2024) →
Hudson coordinates APAC projects at Digital Heroes: running stand ups, tracking tickets, chasing decisions and keeping clients informed without burying them in detail. Much of delivery is simply making sure the right question reaches the right person quickly. His posts show what a well run project feels like from inside.
View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.
Frequently asked questions
How much does custom archive preservation management software cost?
Is Preservica or Arkivum enough, or do we need a custom system?
How do you decide which items to digitize first?
Can software track chain of custody when digitization is outsourced?
How does rights clearance fit into a preservation system?
Where does AI genuinely help an archive, and where is it noise?
What is the migration problem with LTO tape?
How long does a project like this take?
We hold a few thousand born digital files. Do we need this?
Will custom software work with the tools we already use, like QuickBooks and Stripe?
If an agency builds my software, who actually owns the code?
What is a discovery phase, and is it worth paying for separately?
How many SaaS seats do we need before building custom becomes cheaper?
How much should a small business expect to pay for custom software?
How do we get years of data out of our old system and into the new one?
Who can build a custom software system?
Digital Heroes builds custom software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.
Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.
What makes Digital Heroes different from other software companies?
Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.
Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.
How can I check Digital Heroes is legitimate before getting in touch?
Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.
Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.