Media Archive Preservation Software Problems: The 7 That Cost Real Money, and How to Avoid Them
The most expensive failure in an archive build is digitizing material you will never be permitted to show. A vendor batch that comes back clean, checksummed and beautifully catalogued is still wasted money if the music cues on that series were cleared for original transmission only, and the same applies to third party footage on an expired licence and contributors whose consents predate any concept of digital distribution. The spend is unrecoverable, the decay clock did not pause for the carriers you skipped, and the scarce thing you actually burned was vendor capacity and working deck time.
Why does the scope get written as a catalogue instead of a prioritization engine?
Archive briefs almost always describe a catalogue. Records, fields, search, an interface for cataloguers. It is the most familiar shape, most archives already have a collections database, and the request is usually phrased as replacing or improving it.
The question the institution actually needs answered is different: for a given amount of money and time, which items should we do next, and what will we legally be able to do with them afterwards. A catalogue cannot answer that. Prioritization needs three structured inputs multiplied together, decay risk, content value and clearability, and most archives have a partial answer to the second and nothing structured for the first or third.
This is specific to moving image collections because the constraint is not storage cost. It is that the intersection of a playable carrier, a working deck and a person who can operate it is closing on a schedule nobody controls. Acetate keeps hydrolyzing whether or not it is high on your list, magnetic binders keep absorbing moisture, and the pool of engineers who can service certain machines keeps shrinking.
The fix is to specify the prioritization model first and treat the catalogue as an input to it. Ask a developer to whiteboard the data model before you sign: the intellectual work, the manifestation, the physical carrier and the digital instantiation as separate things, because one work can have six carriers of differing quality and three digital copies of differing provenance. A developer who draws assets and files has built a document store and will discover the difference on your programme.
What goes wrong when you migrate decades of catalogue records?
Retrospective data cleanup is the most underestimated line in this category, and it is underestimated because it looks like a data transfer. What is actually there is decades of inconsistent title, series and episode identification, created by different cataloguing conventions under different departmental owners, with several eras of local shorthand embedded in free text fields.
The failure modes are concrete. The same series appears under three titles, one of them a working title that only ever existed internally. Episode numbering restarts mid run because a season was renumbered for a repeat. Shelf locations reference a store that closed in the 1990s. Carrier type sits in a notes field as prose, so the phrase one inch means a tape format in some records and a physical measurement in others. Condition, where it exists at all, is a sentence written by a technician who has retired.
Import that and the prioritization queue inherits every ambiguity, so conservators will not stand behind it, and a ranking nobody will defend in a funding meeting is worth nothing.
The fix is to scope cleanup as its own phase with cataloguers and conservators assigned to it, and to resist the temptation to clean everything. Pick one collection and one carrier family, agree the controlled vocabulary for carrier types and condition fields with the people who will use them, and migrate that slice properly. Leave the rest readable in place. A prioritization engine that is correct for a quarter of the holdings is immediately useful. One that is approximate across all of them is not.
Why do repository and asset management integrations break after launch?
Most archives end up running three systems: a preservation repository underneath for storage, fixity and format management, a media asset management platform for access copies and editorial metadata, and the operational layer that covers condition, prioritization, vendor workflow and rights. The integrations between them fail on ownership of fields rather than on connectivity.
The pattern is predictable. Two systems both believe they hold the authoritative title. A cataloguer corrects a record in one, an editor corrects it differently in the other, and a synchronisation job dutifully propagates whichever ran last. Identifiers drift because one system mints its own on ingest and the other was given an external identifier during a migration nobody documented. Technical metadata extracted at ingest gets written to one system and never reaches the other, so the same file has a frame rate in one place and nothing in the other.
The fix is to write down, before any code, which system owns each field, and to make every other copy explicitly derived and read only. Give every item a single persistent identifier minted in one place and carried everywhere. Then instrument the synchronisation: log every field that changed, in which direction, and surface conflicts as a review queue rather than resolving them silently. Archives think in decades, and a synchronisation that quietly overwrites is a slow corruption of the record.
What happens when per item rights clearance is not covered?
Rights get left out of scope more often than any other requirement here, usually because they feel like a legal matter rather than a software one. The consequence is the failure named at the top of this page: digitized material that cannot be licensed, streamed or exhibited, discovered after the money is spent.
The structure of the problem is that clearance is several independent questions that must all return yes. Music cues cleared for original broadcast only. Performer and contributor consents written before digital distribution existed. Third party footage licensed for a term that expired decades ago. Underlying literary rights. Each has its own evidence, its own territory and its own expiry.
The fix is to make clearance a structured status per item with the evidence attached: the cue sheet, the contract page, the consent form, the territory, the usage type and an expiry that triggers a review. Then weight the digitization queue by clearability, so scarce vendor capacity stops going to material you will never be permitted to distribute. Document extraction earns its place here specifically: cue sheets, contracts and consent forms exist as scanned paper in wildly inconsistent layouts, and a model can pull parties, term, territory and usage type into fields for a researcher to confirm. It does not replace the researcher. It stops the researcher retyping, and it makes ten thousand cue sheets a tractable project rather than an abandoned one.
Should you build custom or configure what you already own?
Some institutions should not build, and the boundary is clear. If your holdings are mostly born digital, in the low thousands of objects, and your obligation is safe custody with a defensible audit trail, Preservica or Arkivum will serve you better than anything commissioned. They are aligned to the reference model for open archival information systems, they are strong on fixity checking, format identification, migration of digital objects and audit trails, and they solve that problem properly. The same applies if you have no funded digitization programme, because operational software for a programme that does not exist is an expensive way to describe an intention.
The pragmatic answer for larger archives is usually not replacement either. Keep the repository for storage, fixity and format management, keep the asset management platform for access copies, and build only the operational layer neither one covers: vault and carrier condition, equipment and vendor scheduling, chain of custody, prioritization and rights. A proposal to replace all of it is quoting a project rather than solving a problem.
Build the operational layer when two or more apply. You hold tens of thousands of physical items across mixed carriers with genuine decay risk. You have multi year funding and must defend the sequencing of that spend to a board or public funder. You use external digitization vendors and cannot reconstruct chain of custody for a batch from last year. Rights uncertainty is blocking monetization of material you already paid to digitize. Or you run more than one vault and items move between them.
How do hidden costs get into the quote?
Carrier variety is the first. Each carrier type needs its own condition schema and its own vocabulary, because acetate film needs an acidity reading, shrinkage measurement and splice condition, while magnetic tape needs binder condition, a note on whether baking was performed and at what temperature, and an assessment of edge damage and pack quality. A quote priced for two carrier families and delivered against nine is a different project.
The second is the data cleanup already described, which is almost always assumed to be background work.
The third is rights document extraction volume. Five hundred scanned cue sheets and ten thousand are not the same engagement, and the difference is not linear because larger corpora contain more layout variants.
The fourth is integration depth with your asset management platform, where the hard part is not the connection but agreeing which system owns which field, and that agreement takes meetings with people who disagree.
What separates a build that works from one that fails here?
Condition capture that a conservator will actually use in the vault decides it. That means format specific fields rather than a free text notes box, tablet entry with barcode scanning at the shelf, and a schema agreed by the people filling it in. A build that stores condition as prose cannot rank anything, and ranking is the entire point.
The second determinant is chain of custody generated by the archive rather than the vendor. The batch manifest is yours, items are scanned out and scanned in, and the vendor posts transfer results, quality control notes, deck used, operator and any interventions through a portal or interface so those facts attach to the item permanently. When a file is questioned in eight years, provenance is the difference between a preservation master and a copy of unknown origin.
Third is the storage plan. Every copy needs a record of where it lives, on what medium, written when, verified when and when it must be migrated. If you use tape, each drive generation reads back only a limited number of previous generations, so every tape written today has a migration horizon whether or not it is budgeted. The one page a board actually wants shows how many items exist as preservation masters, how many copies of each, where they are, when they were last verified and what migration costs over the next five years. Almost no archive can produce that page today without a week of work, and producing it is a reasonable definition of success.
Finally, settle ownership before kickoff: the repository, the cloud accounts and the right to hire anyone else. Archives think in decades, and a system that outlives its original vendor relationship is the only kind worth building here.
The evidence behind this guide
Independent findings on why this investment pays off. Every link goes to the primary source.
- Only 22% of firms are 'future ready' having significantly transformed digitally; these companies show average revenue growth 17.3 percentage points and net margins 14.0 percentage points above their industry average. Source: MIT Center for Information Systems Research (MIT Sloan) (2022) →
- Almost half of all the activities people are paid almost $16 trillion in wages to do in the global economy have the potential to be automated by adapting currently demonstrated technologies. Source: McKinsey Global Institute (2017) →
- An analysis of enrollment and completion data for 221 MOOCs (Katy Jordan, published in the International Review of Research in Open and Distributed Learning, IRRODL, 16(3), 2015 - not the Journal of Distance Education) found completion rates ranging from 0.7% to 52.1%, with a median completion rate of 12.6%, and completion negatively correlated with course length (longer courses had lower completion rates) - underscoring how unsupported self-paced online courses struggle to finish learners. Source: Journal of Distance Education (via ERIC / Katharina Jordan) (2015) →
- Criteo's Global Commerce Review found retail apps convert at 18% versus 4% on mobile web (roughly 4.5x), and travel apps convert at 20% versus 6% on mobile web (about 3.3x). Source: Criteo (2017) →
Prasun founded Digital Heroes in 2017 and leads it from New York. His work sits where commercial decisions meet delivery: which projects to take on, how teams are shaped across five offices, and where a build is likely to go wrong. Readers get the view from the side that owns the outcome.
View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.
Frequently asked questions
Why is prioritization harder than it looks in a moving image archive?
Because a defensible ranking needs three structured inputs multiplied together, decay risk from format specific condition data, content value and rights clearability, and most archives have only a partial answer to the second. Without condition captured as structured fields, acidity readings, shrinkage, binder condition, you cannot rank at all, and a queue built on title recognition is impossible to defend to a funder when it puts a famous drama ahead of a rotting regional news series.
What makes catalogue migration so slow in archives?
Decades of inconsistent title, series and episode identification created under different cataloguing conventions, with carrier type and condition buried in free text written by staff who have retired. Working titles that only existed internally, renumbered seasons and shelf references to stores that closed all survive in the data. Scope cleanup as its own phase with cataloguers and conservators assigned, and migrate one collection and one carrier family properly rather than approximating across everything.
How do integrations with a preservation repository actually fail?
On field ownership rather than connectivity. Two systems both believe they hold the authoritative title, identifiers drift because one mints its own on ingest, and a synchronisation job propagates whichever record was edited last. Nothing goes down, so the failure appears as staff quietly learning which system to trust for which field. Write down which system owns each field before any code, and surface synchronisation conflicts as a review queue rather than resolving them silently.
What is the cost of leaving rights clearance out of the first release?
Digitizing material you will never be permitted to show, which is unrecoverable spend plus the vendor capacity and working deck time you did not use on carriers that were actually dying. Clearance is several independent questions that must all return yes: music cues, performer and contributor consents, third party footage terms and underlying rights, each with its own evidence and expiry. Make it a structured status per item and weight the queue by clearability.
Where does document extraction genuinely help an archive?
On scanned cue sheets, contracts and consent forms, pulling parties, term, territory and usage type into structured fields for a researcher to confirm. That is what turns ten thousand paper documents from an abandoned project into a tractable one. What does not work is asking a model to judge preservation priority or rights status on its own, because both need evidence you can show a funder or a lawyer rather than a confident answer.
Should we replace Preservica or Arkivum, or build around them?
Build around them in almost every case. They handle storage, fixity, format identification and audit trails properly, and reimplementing that is wasted budget. What they do not run is your vault, carrier condition, scarce playback equipment scheduling, vendor batches and chain of custody, or per item rights. A proposal to replace the whole stack is quoting a project rather than addressing the part that is actually hurting.
Which cost is most often missing from an archive software quote?
Vocabulary and carrier variety. Every carrier type needs its own condition schema and its own agreed terminology, and conservators, cataloguers and the rights team frequently use different words for the same thing. Archives with a documented condition survey method move noticeably faster, and if you do not have one, that definition work belongs in the plan with an owner and a date rather than being absorbed during the build.
How do we know the build has succeeded?
Two tests. Conservators enter condition data at the shelf without being chased, because the fields match how they actually assess carriers. And you can produce, on demand, a single page showing how many items exist as preservation masters, how many copies of each, where they live, when each was last verified and what migration will cost over five years. Most archives need a week of manual work for that page today.
What does a $50,000 custom software budget actually buy?
Is a solo freelancer enough for my project, or do I really need an agency?
We run everything on spreadsheets and Airtable. How do we know it's time for custom software?
How many people should be working on my software project?
How many SaaS seats do we need before building custom becomes cheaper?
What should I have ready before I contact a development agency?
We run everything on Airtable and spreadsheets. When is it time to go custom?
Who can build a custom software system?
Digital Heroes builds custom software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.
Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.
What makes Digital Heroes different from other software companies?
Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.
Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.
How can I check Digital Heroes is legitimate before getting in touch?
Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.
Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.