Clinical Trial Site Software: Fixing the Paper Problem
Build if you run three or more sites, or one site running 15+ concurrent trials, and your coordinators are still keeping enrollment on a whiteboard and source on paper. In Digital Heroes delivery experience a focused first release costs $60k to $130k and ships in 12 to 16 weeks; a full site platform runs $150k to $400k phased over 6 to 12 months. Below that scale, a mid-tier CTMS plus disciplined paper is defensible. Above it, the money you lose to missed visit windows, screen failures nobody forecasted, and monitors billing you for source verification chaos will pay for the build inside two years.
Why site software makes or breaks a research site operator
Walk into most independent research sites at 7:40am and you will find the same three artifacts. A dry-erase board with columns for each active protocol and magnet dots for subjects in screening. A shared Outlook calendar where visit windows live as all-day events with the protocol number typed into the subject line. And a wall of three-ring binders, one per subject, holding source worksheets that a coordinator prints from a Word template the sponsor emailed in a zip file eight months ago.
The sponsor gave you their EDC. Medidata Rave, Veeva CDMS, Oracle Clinical One, maybe Medrio for a smaller study. Those systems belong to the sponsor and only accept clean data after it exists. They do not tell you which of your 140 active subjects is drifting toward a protocol deviation on Thursday. They do not know your coordinator ratio. And they are per-study, so a site running 22 concurrent protocols is logging into nine different portals with nine password policies. Your CTMS, if you have one, is probably Clinical Conductor or RealTime CTMS, and it holds the financial contract data reasonably well while ignoring the thing that actually kills you: the daily operational math of who is due, who is late, and what has not been signed.
The leak is specific. A coordinator at a three-site cardiology network runs a Day 90 visit, completes the source worksheet on paper, and the binder goes to the file room. Nineteen days later the CRA arrives for monitoring, opens the binder, and finds the vitals were recorded but the concomitant medication log was never updated from the visit note. That is a query, then a deviation, then a CAPA, then an hour of the PI's time signing a memo to file. Multiply by four monitoring visits a year across 22 protocols. Across the sites we have built for, coordinators lose 6 to 9 hours a week to transcription, chasing signatures, and reconciling paper against the EDC, and that is before the day a sponsor amendment changes the visit schedule and someone has to hand-recalculate 60 windows.
Problem 1: Visit windows are calculated by humans and humans miss them
Protocol says Day 90, window minus 7 plus 7, anchored on first dose. Subject 014 dosed on a Tuesday. The coordinator does the arithmetic in her head, puts an Outlook event on the wrong anchor date because she counted from randomization instead of first dose, and the subject shows up on Day 99. Out of window. Deviation. If that subject is in the primary endpoint population and the sponsor is strict, you have just donated a data point.
Off-the-shelf CTMS products do carry visit schedules, but they are built for the version of the protocol that existed at study startup. When Amendment 3 arrives and moves the Week 12 window from plus or minus 7 to plus or minus 3 and adds an unscheduled PK draw, most sites re-enter the schedule by hand or, more honestly, do not re-enter it at all and rely on the lead coordinator remembering. The tool has no concept of a versioned schedule of assessments that applies from a given date forward for subjects who re-consent.
A custom build models this properly. The core object is a protocol version with a schedule of assessments: each visit has an anchor event (first dose, randomization, prior visit), a day offset, a window in days, and a list of required procedures with the equipment and staff role each needs. Subjects carry a consent version pointer. When Amendment 3 lands, you create version 3, mark which subjects re-consented and when, and the engine recomputes forward windows only for those subjects while leaving the rest on version 2. The coordinator dashboard sorts by days-to-window-close, not by date, so the thing closing Thursday is at the top regardless of which of your 22 protocols it belongs to. Every recomputation is logged with who, when, and why, because your monitor will ask.
Problem 2: Source data lives on paper and monitoring visits become archaeology
The paper source worksheet is not stupid. It is fast, it works when the network is down, and coordinators trust it. What it cannot do is tell you at 4pm on a Friday which of the week's 31 visits have incomplete source before the CRA lands Monday. So instead the coordinator spends Thursday and Friday pulling binders and flipping pages, and the CRA spends the first four hours of a two-day visit doing source data verification by eye.
Sponsor EDC does not solve this because it is downstream. eSource products exist, but the generic ones make you rebuild every worksheet as a form from scratch, which is a two-week job per protocol and is why sites abandon them by protocol four.
What we build instead: the worksheet template is the schema. You upload the sponsor's Word source worksheet and an extraction step, using a document model, reads the field labels, units, and normal ranges into a draft form definition that a coordinator corrects in about 20 minutes rather than rebuilding in two weeks. This is the one place AI pays for itself cleanly in this category, because the input is a structured document, the output is reviewable before it ever touches subject data, and a human approves the schema. From there the form runs on a tablet at the visit, fields are typed with units and ranges so an out-of-range hemoglobin flags at entry rather than at query, and every save writes an append-only audit row with user, timestamp, old value, new value, and reason for change. Part 11 compliant electronic signature on the PI review step. Then a monitoring packet view: filter to one CRA's assigned subjects, show every field changed since last visit, export a read-only certified copy. Sites we have built this for cut source verification prep from two days to under an hour.
Problem 3: Enrollment forecasting is a guess and the guess costs you the next contract
The sponsor asks for your enrollment projection during feasibility. You look at the whiteboard, remember that the last cardiology study took eleven months to hit 40, and say "we can do 25 in six months." You commit. You get 9. Now you are a low enroller in that sponsor's site selection database and the next three studies go elsewhere.
Nothing off the shelf helps here because your CTMS knows contracted enrollment but not your actual funnel: how many charts you pre-screened, how many passed I/E on paper, how many consented, how many screen-failed and on which specific criterion. That data exists only in a coordinator's pre-screening log, which is a spreadsheet on one laptop.
The build makes the funnel a first-class object. Every referral source is tracked (EHR query, physician referral, patient registry, radio, community event), every pre-screen records the specific I/E criterion that killed it, and screen failure reasons are coded not free text. Once you have twelve months of that, forecasting becomes arithmetic instead of vibes: this protocol's I/E overlaps 71 percent with a study you ran in 2024, that study converted 6.2 percent of pre-screens to randomized at a rate of 34 pre-screens per month from your EHR cohort, so your honest projection is 14 in six months and you should either negotiate the criteria or decline. Sites that walk into feasibility with a funnel history instead of an anecdote win better contracts, because sponsors can tell the difference. The EHR side is where you plug in a cohort query against Epic via FHIR or a nightly extract, so the pre-screen list is generated rather than hand-built.
Problem 4: Coordinator capacity is invisible until someone quits
A site director cannot tell you today how many hours of protocol-mandated work sit on each coordinator next week. So allocation happens by whoever complains loudest, the strongest coordinator absorbs the hardest protocol, burns out, and leaves after fourteen months taking the institutional memory of six studies with her. At the sites we work with, replacing a certified CRC and getting them protocol-trained lands somewhere around $40k to $60k once you count recruiting, ramp, and the deviations that happen during ramp.
CTMS tools track staff assignment as a name field. They do not know that a Day 1 dosing visit on an oncology protocol is 5.5 hours of coordinator time with a two-hour PK series and the Day 30 follow-up on a device study is 40 minutes.
Custom fixes this because you already modeled procedures per visit in Problem 1. Attach a time estimate and a required role to each procedure, and next week's schedule becomes a load chart per coordinator in hours, not a list of names. Now you can see that Maria is at 47 booked hours before any unscheduled visits and that the Thursday PK draw collides with the only centrifuge. It also feeds the thing your CFO wants: coordinator hours per visit against the per-visit payment in the clinical trial agreement, which tells you which protocols are actually profitable. We have watched a site discover a study paying $1,850 per visit was consuming $2,100 of loaded coordinator and PI time. They stopped bidding on that sponsor's device studies.
Problem 5: Subject retention runs on one coordinator's phone
Retention is the whole business. A subject who drops at Week 24 of a 52-week study is a data point you already paid for and cannot bill. Sites reduce dropout with reminder calls, and those calls happen when a coordinator has a gap, which is never, so they happen on Fridays or not at all.
Generic scheduling tools cannot do this because the reminder cadence is protocol-specific and the confirmation has to be logged as a contact attempt for the monitor.
Build: an outbound reminder engine that fires against the computed window, not a calendar entry, so a rescheduled visit moves its own reminders. SMS at T-7, T-2, T-1 with the subject's preferred language and the visit-specific prep instructions (fasting, hold morning dose, bring diary). Two-way, so a reply of "cannot make Tuesday" opens a reschedule slot inside the remaining window rather than outside it. Where an AI voice agent actually helps: after-hours inbound. Subjects call at 8pm because that is when they remember. A voice agent that can confirm a visit, capture a reschedule request into the window engine, and hard-escalate to the on-call coordinator the instant anything sounds like an adverse event, with a full transcript logged. Never let it triage symptoms. It confirms, reschedules, and escalates, and it logs every contact attempt so the retention record is complete for the monitor.
What this costs at a research site
Across 2,000-plus projects, Digital Heroes typically ships a focused first release at $60k to $130k in 12 to 16 weeks. For a research site that first release is almost always: protocol and schedule-of-assessments model with versioning, subject roster and consent tracking, the window engine plus coordinator dashboard, and eSource with audit trail and Part 11 signatures for two to three protocols. That is the release that stops the bleeding.
Full platforms, meaning the funnel and forecasting, capacity model, CTA-linked financials with per-visit invoicing, EHR cohort integration, sponsor EDC reconciliation, and the retention engine, run $150k to $400k phased over 6 to 12 months.
What drives price up specifically in this category: 21 CFR Part 11 validation work is real and it is not optional, so budget for IQ/OQ/PQ documentation, a validation plan, and a traceability matrix, which in our delivery experience adds 15 to 25 percent to the engineering line. Multiply that if you take EU subjects and inherit Annex 11 and GDPR. Epic or Cerner integration is the second driver: a read-only FHIR cohort query through a vendor's app orchard process is months of paperwork before a line of code, and a nightly flat-file extract from your health system's IT group is often faster and cheaper, so decide early. Third: the number of sponsor EDCs you want to push to. One EDC integration is scoped work; five is a platform problem, and most sites should not build EDC push at all in v1 because the sponsors change systems per study and you will chase them forever.
Build versus buy: take the position
Buy if you are a single site running under about eight concurrent protocols with two or three coordinators. Clinical Conductor or RealTime CTMS plus disciplined paper source will hold you, and the subscription is cheaper than any build. Buy also if you are a site network that has genuinely standardized on one therapeutic area and one or two sponsors, because then the sponsor's own tooling covers more of your surface than it would otherwise.
Build when these signals show up, and they show up together. You are running three or more sites, or one site past 15 concurrent protocols. Your deviation count is climbing and more than a third of deviations are visit-window or missing-source, which means the problem is operational, not clinical. You have hired a person whose job is essentially to be a human integration layer between the whiteboard, the CTMS, and nine EDC portals. You lost a study to a sponsor because you could not evidence your enrollment history. And the tell that ends the argument: your best coordinator has a personal spreadsheet that the entire site depends on. That spreadsheet is your requirements document, and the fact that it exists means the off-the-shelf tool has already failed. Buy the CTMS for the money and the contracts; build the operational layer that the CTMS refuses to be.
How to choose a developer for clinical trial site software
Ask them to model a schedule of assessments on a whiteboard, cold, in fifteen minutes. If they do not immediately ask what the anchor event is, whether windows are calendar or business days, and how re-consent under an amendment affects forward visits, they have never built this and will discover those requirements in month four at your expense.
Ask what they have shipped under 21 CFR Part 11. Specifically: have they produced a validation plan and traceability matrix, and have they built an append-only audit trail with reason-for-change capture, and have they sat through a sponsor audit of a system they built. Part 11 is not a checkbox a developer adds at the end; it shapes your data model from day one, because you cannot retrofit immutability onto a table that has been doing UPDATEs.
Ask about their integration scars. The honest answer to "can you integrate with Epic" is a question back: read-only or write, FHIR or extract, and has your health system's IT group approved app orchard access, because that timeline is theirs, not ours. A developer who says yes without asking is going to blow your schedule.
Finally: get the code ownership and the source escrow in writing before kickoff, and make the first release genuinely narrow. The window engine and eSource for three protocols, live in 14 weeks, in coordinators' hands, is worth more than a beautiful full platform that arrives after your best CRC has already quit.
The evidence behind this guide
Independent findings on why this investment pays off. Every link goes to the primary source.
- The average developer spends more than 17 hours a week dealing with maintenance issues such as debugging and refactoring, and about four of those hours on 'bad code' - waste that equates to nearly $85 billion annually worldwide in opportunity cost. Source: Stripe (2018) →
- Per the Standish Group CHAOS 2020 report (reviewed at this URL), across tens of thousands of software projects roughly 31% end successfully, about 50% are 'challenged', and roughly 19% fail outright; small projects succeed far more often than large ones, and Agile approaches succeed at markedly higher rates than Waterfall. Source: The Standish Group (2020) →
- A study (led by Prof. Pak-Lok Poon, published in Frontiers of Computer Science, 2024) reviewing decades of spreadsheet-quality research found that about 94% of spreadsheets used in business decision-making contain errors, illustrating the hidden risk of manual spreadsheet workarounds that custom software is built to replace. Source: Central Queensland University / phys.org (Prof. Pak-Lok Poon et al.) (2024) →
- In an October 2025 survey of 530 small-business employers (conducted by TechnoMetrica, October 3-9, 2025), 88% reported using AI tools and 73% said those tools had been important to their competitiveness and growth over the past year, with 60% citing efficiency and productivity as the primary motivation for adoption (42% cited improving customer service). Source: Small Business & Entrepreneurship Council (SBE Council) (2025) →
Rohan advises mid-market and enterprise teams on ERP, CRM and custom software, and has led delivery on dozens of business-software builds.
Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.