Industry guide · Custom Software

eTMF Software for Sponsors and CROs: Why Your Completeness Number Lies

Etmf Management software visual showing folder tree, task checklist, and compliance shield.
The short answer

Expect $95,000 to $190,000 and 14 to 20 weeks for a first eTMF release covering a modified reference model index, milestone-driven expectedness, site and vendor ingestion and a completeness view your TMF lead will actually defend. A validated platform adding legacy migration, per-country redaction, partner access and an inspection export runs $260,000 to $650,000 phased over 9 to 15 months. Build when you run a modified index across more than roughly a dozen concurrent studies, when you are a CRO carrying TMFs for several sponsors, or when acquisitions left you with three legacy systems. Do not build if you run one or two studies and have no quality function to own computerised system validation. Licence Veeva Vault eTMF and hire a TMF manager instead.

Why a trial master file breaks the software you buy for it

An inspection notice arrives for the study your filing depends on. The eTMF dashboard says 94 percent complete. Then someone opens the country file for Poland and finds a delegation of authority log signed by an investigator who left in month four, three versions of the laboratory manual with nothing recording which one superseded which, and monitoring visit reports for visits six through nine still sitting in a departed CRA's mailbox. That 94 percent was never a measure of completeness. It was a count of documents that happened to be filed, divided by a list somebody typed into a configuration screen at study start and never revisited.

ICH E6 expects a trial master file that permits evaluation of trial conduct and of the quality of the data produced. An inspector reads it as the story of how the trial was run. That story is told by which artifacts exist, when they were created relative to the events they document, and how they attach to a site, a country, a vendor and a milestone. A document repository with a folder tree holds pages. It does not hold the story, and the gap between the two is filled by a TMF team doing manual reconciliation for six weeks before every inspection.

Problem 1: expectedness is a calculation, not a folder structure

The DIA TMF Reference Model gives you zones, sections and artifacts, and almost every sponsor and CRO has modified it. The modification is not the hard part. The hard part is expectedness: which artifacts should exist right now, for this study, in this country, at this site, given where each of them sits in its own lifecycle. A site that has not activated should not be flagged for a missing monitoring visit report. A site that activated eleven weeks ago with no visit report on file is a finding in waiting. One number cannot represent both.

Veeva Vault eTMF models this properly and is the strongest product in the category. What sponsors run into is that expectedness rules are configuration inside a platform, so when your rules depend on facts the platform does not hold, for example a vendor's contracted scope, a local ethics requirement in one country, or a device study's technical documentation, the real truth migrates back into a spreadsheet sitting beside the system that was meant to hold it. Montrium eTMF Connect inherits the content model of SharePoint and Microsoft 365, which is comfortable for IT and awkward once a single study carries tens of thousands of artifacts with long version chains. Phlexglobal PhlexTMF comes from a TMF services heritage, which is genuinely useful if you want people to run the TMF for you and less useful if you want to own and change the rules yourself. Florence eBinders is excellent on the site side, where the investigator site file and remote monitoring access are the real problems, but it is not a sponsor-side portfolio view across studies. MasterControl treats TMF content as controlled quality documents, which is a different shape from study, country, site and vendor.

What a custom build does: it makes expectedness a rule engine sitting on a milestone timeline. Each artifact definition carries a trigger, an owner, a due offset and a scope, so the system generates the expected set continuously rather than once. Activate a site in Spain and 31 expected artifacts appear against that site with dates. Terminate a vendor and their expected set closes rather than sitting open forever, dragging your number down. This one change is usually what converts the TMF from a filing cabinet into an operational instrument, because now the weekly report tells a study manager what to chase this week instead of telling everyone the same 94 percent.

Problem 2: documents arrive from everyone except the people you employ

Sites email scans. Central labs post manuals to their own portal. Ethics committees send approvals as photographed paper in three languages. Couriers, translation vendors, IRT providers and imaging core labs each have a different transfer habit. Your CRAs upload from the field on a hotel connection. The result is duplicates, wrong site numbers, wrong study, and a filing queue that a coordinator works through by opening each PDF and reading it.

This is where AI does one concrete job, and it is not a chatbot. A classification and extraction pass reads the incoming document, proposes the artifact type against your index, pulls the site number, the document date, the version and whether a wet or electronic signature is present, then routes it to a human for one-click confirm. Across our regulated document projects, the no-touch confirmation rate settles around 80 to 90 percent after a few weeks of corrections, and the remaining 10 to 20 percent is exactly the interesting material a reviewer should be looking at anyway. A second model pass detects near-duplicates and probable supersedes, which is the failure mode that produces three laboratory manuals with no relationship between them. Every automated decision has to be recorded in the audit trail as a system action with the model version, because an inspector will ask who classified this and the honest answer needs to be reconstructable.

Problem 3: migration is the project, and everyone budgets it as a task

Legacy paper in offsite storage. A shared drive with 400,000 files. The previous CRO's system, which you can export from but not query. An acquired asset whose TMF lives in a third platform under a licence that ends in four months. Every eTMF programme we have seen run over budget ran over on migration, not on features.

The approach that works is to inventory before you decide anything. Crawl the sources, hash and cluster, classify with the same model that will run in production, then produce a report showing what you actually have against what the reference model expects for those studies. That report is the artefact your quality lead uses to decide what gets migrated, what gets archived as-is with a documented rationale, and what is not TMF content at all, which is usually a surprising share of a shared drive. Migrating blind means paying to carry rubbish into a validated system, then paying again to explain it to an inspector.

Problem 4: you need three numbers, not one

Completeness is filed against expected. Timeliness is the gap between the document date and the filing date measured against your own SOP, and it is the number inspectors probe hardest, because a file that was assembled the month before an inspection tells its own story. Quality is the QC pass rate on reviewed artifacts, which only means anything if your QC sampling plan is defined and applied consistently. Reporting one blended percentage hides all three. Build the model so any of them can be sliced by study, country, site, vendor and TMF owner, and so the same figures can be regenerated for a date in the past, because the question at inspection is what the TMF looked like then, not what it looks like today.

Problem 5: privacy, blinding and who is allowed to see what

Consent forms arrive with patient initials and dates of birth on them. GDPR obligations differ from the redaction habit at a US site. A blinded study needs randomisation and unblinded pharmacy content walled off from the study team but retained. A CRO holds TMFs for competing sponsors in the same instance and cannot let one see the other. An inspector needs read access scoped to one study with every view logged. Access control here is not a role dropdown, it is a policy layer over study, country, site, artifact type and blinding status, and it has to be testable, because you will be asked to demonstrate it.

Validation is not the line item to cut

Anything holding TMF records is a regulated computerised system: 21 CFR Part 11 in the US, EU Annex 11 in Europe, and GAMP 5 as the practical framework. That means a validation plan, requirements traced to test scripts, installation, operational and performance qualification, documented change control and a periodic review. In our experience validation and its documentation add 15 to 25 percent on top of engineering cost, and a partner who quotes a regulated eTMF without that line has either done it before and hidden it, or has never done it. Ask which.

What this costs and how long it takes

A first release with your modified index, milestone-driven expectedness, ingestion including the classification pass, QC workflow and the three metric views runs $95,000 to $190,000 and ships in 14 to 20 weeks, based on Digital Heroes delivery experience. A validated platform adding legacy migration, per-country redaction, sponsor and vendor portals, e-signature and a defensible inspection export runs $260,000 to $650,000 phased over 9 to 15 months.

What drives the number up in this category specifically: the volume and mess of legacy content, since a million-file shared drive is a different project from a clean export. Multi-tenant separation if you are a CRO holding several sponsors. Electronic signature, because Part 11 signature manifestations have to be right rather than approximately right. The number of external feeds, since each vendor transfer is its own contract, format and failure mode. And translation handling, if your ethics documents arrive in nine languages and someone has to certify the English version. What keeps it down: scoping the first release to active studies only, leaving closed studies in their archive with a documented plan, and resisting the urge to rebuild the CTMS at the same time.

Build versus buy, and when licensing is the right call

Licence, and do not call us, if you are a small sponsor with one or two studies, no in-house quality function and no appetite to own validation. Veeva Vault eTMF or Florence will serve you better than anything you commission, and the per-study cost is rational at that volume. The same is true if your TMF is genuinely standard and your operations team is happy to work the way the product works.

Build when two or more of these hold. You run a materially modified index and are maintaining the real expectedness rules outside the system. You are a CRO and need one portfolio view across sponsors with hard separation. You are past roughly a dozen concurrent studies and per-study platform fees have overtaken a build. You have vendor and site feeds that nobody will change for you. Or you have TMFs in three systems after an acquisition and consolidation is now a board-level commitment with a date on it. The tipping point is not features, it is that your expectedness logic has become a piece of intellectual property and you cannot keep it in a spreadsheet.

How to choose a developer for a regulated eTMF

Ask them to whiteboard expectedness before you sign anything. Someone who has done this will draw artifact definitions, triggers, scope by country and site, and a lifecycle that opens and closes expected sets. Someone who draws documents and folders has built a file store and is about to learn ICH E6 on your budget.

Ask what validation deliverables they produce and who writes them. Requirements traceability matrix, IQ, OQ, PQ, change control and a periodic review procedure should come back without hesitation. Ask how they handle supersede chains and how the system reconstructs the TMF as of a past date, because that single question separates people who have faced an inspection from people who have read about one.

Ask who owns the code, the infrastructure accounts and the validation package, and get it in writing before kickoff. The validation documentation matters as much as the repository, because without it your next partner revalidates from zero. At Digital Heroes the client owns both from the first commit, and we would tell you to walk away from anyone who hedges on it.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. The right combination of digital transformation actions can unlock as much as US$1.25 trillion in additional market capitalization across Fortune 500 companies, while the wrong combinations put more than US$1.5 trillion at risk; companies with all three core factors (strategy, aligned technology, and change capability) saw a 5% market-value lift relative to peers. Source: Deloitte (2023) →
  2. Deloitte's research found that digitally advanced small businesses experienced revenue growth nearly 4x as high as the prior year, were about 3x as likely to have exported, were nearly 3x as likely to have created new jobs, and were more than 3x as likely to have seen more sales inquiries in the last year. Source: Deloitte (research summarized by Google) (2017) →
  3. In the Flexera 2025 State of ITAM report, respondents reported roughly 33% of SaaS spend is wasted, underscoring how paying for off-the-shelf seats and tiers that go unused erodes the supposed cost advantage of generic SaaS. Source: Flexera (2025) →
  4. An analysis of enrollment and completion data for 221 MOOCs (Katy Jordan, published in the International Review of Research in Open and Distributed Learning, IRRODL, 16(3), 2015 - not the Journal of Distance Education) found completion rates ranging from 0.7% to 52.1%, with a median completion rate of 12.6%, and completion negatively correlated with course length (longer courses had lower completion rates) - underscoring how unsupported self-paced online courses struggle to finish learners. Source: Journal of Distance Education (via ERIC / Katharina Jordan) (2015) →
Indi W. · Mobile Designer · Sydney

Indi designs mobile app screens at Digital Heroes, working through the states an interface needs before it can be built: loading, empty, error, success. It is detailed work that decides how an app feels in the hand. Useful reading if you are scoping an app and wondering where design hours go.

View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.

FAQ

Frequently asked questions

How much does it cost to build a custom eTMF system for a sponsor or CRO?
A first release covering a modified reference model index, milestone-driven expectedness, document ingestion and completeness reporting runs $95,000 to $190,000 and ships in 14 to 20 weeks, based on Digital Heroes delivery experience. A fully validated platform with legacy migration, redaction, partner portals and inspection export runs $260,000 to $650,000 phased over 9 to 15 months. Budget another 15 to 25 percent on top of engineering for computerised system validation documentation. Legacy content volume is the single biggest cost driver, not features.
Is Veeva Vault eTMF good enough, or should we build our own?
Veeva Vault eTMF is the strongest product in the category and is the right answer for a sponsor running a handful of studies with a standard index and no in-house appetite to own validation. It becomes a poor fit when your expectedness rules depend on facts the platform does not hold, such as vendor contracted scope or local ethics requirements, because the real logic then lives in a spreadsheet beside the system. Building makes sense once you are running many concurrent studies on a materially modified index, or you are a CRO needing one portfolio view across sponsors.
What does TMF completeness actually mean, and why do dashboards overstate it?
Completeness is filed artifacts divided by expected artifacts, and dashboards overstate it because the expected list is usually typed in once at study start and never recalculated as sites activate, countries open and vendors change scope. A site that activated eleven weeks ago with no monitoring visit report on file should be pulling your number down, and often is not. You also need timeliness, the gap between document date and filing date, and QC pass rate, because completeness alone hides a file assembled the month before an inspection.
How long does an eTMF migration from shared drives and a legacy system take?
Migration is where eTMF programmes overrun, so plan it as its own workstream rather than a task. The pattern that works is to crawl and classify the sources first, produce an inventory against what the reference model expects, and let your quality lead decide what migrates, what is archived as-is with a rationale and what is not TMF content at all. For a mid-sized portfolio, expect the migration workstream to run in parallel across several months rather than weeks.
Does a custom eTMF need 21 CFR Part 11 validation?
Yes. Any system holding trial master file records is a regulated computerised system, so 21 CFR Part 11 in the US and EU Annex 11 in Europe apply, with GAMP 5 as the practical framework. That means a validation plan, requirements traced to executed test scripts, installation, operational and performance qualification, documented change control and periodic review. Expect this to add roughly 15 to 25 percent on top of engineering cost, and treat any quote that omits it as incomplete.
Can AI classify and file TMF documents automatically?
It can propose classification and metadata, and that is where the value is. A model reads the incoming scan, suggests the artifact type against your index, extracts the site number, document date and version, and flags probable duplicates and supersedes, then a human confirms with one click. No-touch confirmation typically settles around 80 to 90 percent after a few weeks of corrections. Every automated decision must be written to the audit trail as a system action with the model version, because an inspector will ask who classified the document.
How do you keep a CRO's eTMF separated between competing sponsors?
Separation has to be a policy layer over study, sponsor, country, site, artifact type and blinding status, not a role dropdown, and it must be demonstrable in a test script rather than asserted. In practice that means every query is scoped at the data layer, unblinded content is walled off from study team roles while remaining retained, and access is logged at the view level. If you are a CRO, raise this in the first design session, because retrofitting tenancy after a system is validated is expensive.
What should we ask a development partner before commissioning an eTMF build?
Ask them to whiteboard the expectedness model: artifact definitions, triggers, scope by country and site, and lifecycles that open and close expected sets. Ask which validation deliverables they write and whether they have produced a traceability matrix and executed qualification scripts before. Ask how the system reconstructs the TMF as of a past date, because that question separates teams who have supported an inspection from teams who have not. Finally, confirm in writing that you own the code, the infrastructure accounts and the validation package.
We have TMFs in three different systems after an acquisition. Is consolidation worth building?
This is one of the clearest build cases in the category, because none of the three vendors will model the others' indexes and you will otherwise carry three licences, three sets of SOPs and three inspection stories indefinitely. The work splits into an inventory and classification pass across all sources, a mapping to one target index, and a decision log recording what was migrated and what was archived in place. Expect the consolidation to be dominated by content decisions your quality team must make, not by engineering.
Should we build an MVP first or go straight to the full system?
MVP first, for almost everyone: ship the single workflow that carries the business value in 10 to 16 weeks, learn from real users, then fund phase two from evidence instead of guesses. The caveat is that an MVP is a small version of a well-built system, not a badly built version of a big one; the data model must already support what comes next. An agency that cannot tell you what they deliberately left out of your MVP has not designed one.
Our developer disappeared mid-project. Can another team pick up the code?
Yes, this is a routine engagement, provided the code exists somewhere you can access, so your first move is securing the repository, hosting, and domain credentials today. A takeover starts with a one to two week paid code audit that ends in one of three verdicts: continue the build, keep the design but rebuild the weak parts, or start over. Digital Heroes has inherited enough projects to say plainly that sometimes the rebuild is cheaper than the rescue, and an honest agency will tell you which one you have before taking your money.
Who owns the code when an agency builds my software?
You should, completely, through a written intellectual property assignment that transfers everything on final payment; without that clause, copyright stays with whoever wrote the code by default. Insist that the repository lives in your own GitHub organization from day one and that hosting, domains, and third-party accounts are registered to you. Also check for licenses to the agency's proprietary frameworks buried in the contract, because those can make switching vendors practically impossible even when you own your own code.
What should I have ready before I contact a development agency?
Three things, none of them technical: a one-page description of the problem in your own words, a list of the tools and spreadsheets the new system must replace or connect to, and a must-have versus nice-to-have split of features. Add a budget range, even a wide one, because it changes the conversation from fantasy to engineering. You do not need a formal specification; producing that is what a discovery phase is for.
Can I build my product on a no-code tool like Bubble instead of hiring developers?
For testing whether anyone wants the product, yes, and Bubble's paid plans start at $29 a month, which is the cheapest validation you will ever buy. The ceiling arrives with complex data relationships, heavy integrations, performance at a few thousand users, and the fact that you cannot export a Bubble app to servers you control. A path many Digital Heroes clients take: prove demand on no-code, then rebuild custom once revenue justifies it, treating the no-code version as a paid prototype rather than a foundation.
How many SaaS seats do we need before building custom becomes cheaper?
The crossover usually shows up between 20 and 50 seats on premium tiers. Salesforce Enterprise lists at $165 per user per month, so 40 users cost about $79,000 a year in subscriptions, which is real money against a custom system you would own outright. Run the comparison over three years: if subscription spend beats the build cost plus 15-20% annual maintenance, custom wins on price before you even count workflow fit.
If we build for 20 users now, will the software cope with 500 later?
It should, without a rewrite, if it was built on a standard cloud stack; going from 20 to 500 users is mostly a hosting configuration change costing hundreds a month, not a second project. What actually breaks under growth is sloppier work: database queries never indexed for volume and features designed assuming one office's worth of data. Before signing, ask the vendor what happens to the system at ten times today's data, and listen for a specific answer.
How do I calculate whether custom software will pay for itself?
Divide the build cost by the monthly benefit, where benefit is hours saved times loaded hourly cost, plus subscription fees replaced, plus any revenue the software unlocks. Three staff saving 10 hours a week each at a $40 loaded rate is about $62,000 a year, which pays back a $60,000 build in roughly 12 months. Across Digital Heroes internal-tool projects, 12 to 24 months is the normal payback range, and anything projecting under 6 months usually means the spreadsheet is hiding costs.
Who can build a custom software system?

Digital Heroes builds custom software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other software companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading
let's build

Build something worth launching.

A plan, a team, a timeline, within 24 hours. No decks, no discovery calls. Tell us what you're building and we'll come back with a real scope and a real number.

message us directly · we reply within one business day

mission briefing

Monthly dispatch

Playbooks, real build costs, and what we're shipping. One email a month. No fluff.

visit us

New York HQ

1140 Broadway, Suite 704 · New York, NY 10001

Get directions
Online now

Hey there 👋 How can we help you today?