Industry guide · Internal Tools

CDISC Submission Standards Software: Why Every Legacy Study Costs You Six Weeks of Remapping

Cdisc Submission Data Standards software visual showing table properties, git compare arrows, and approved record.
The short answer

Expect $85,000 to $170,000 and 12 to 18 weeks for a first release covering a metadata repository, a reusable raw-to-SDTM mapping library, generated Define-XML and an automated validation loop against published conformance rules. A full platform adding ADaM derivation traceability, sponsor controlled terminology governance, reviewer guide generation, legacy study onboarding and regeneration when a standard version changes runs $220,000 to $550,000 phased over 8 to 14 months. Build if you file more than a couple of submissions a year or carry acquired assets in unfamiliar structures. If you run one study and one submission, licence Pinnacle 21 Enterprise and hire a good contract programmer.

Why standards work eats a programming department alive

A submission is eleven weeks out. The statistical programming lead is looking at a study acquired with a portfolio two years ago, where raw data arrived from a CRO in a structure that resembles nothing else you own. The mapping specification is a spreadsheet written by someone who has left. Pinnacle 21 returns several hundred findings, most of which are the same three problems repeated across domains, and each fix means editing a mapping program, rerunning, regenerating the Define-XML by hand because the last person did that by hand, and re-reading the reviewer guide to see whether the narrative still matches. Six weeks disappears into a loop that produces no science.

Every sponsor above a certain size has this loop. It is not caused by CDISC being difficult. It is caused by mapping knowledge being stored as programs rather than as metadata, so nothing learned on study seven is available to study eight, and a change in the standard version means opening every program again.

Problem 1: your mapping library is a folder of programs, not an asset

The same raw structures recur across your portfolio. Your CRFs are mostly standard. Your laboratory vendors send the same layouts. Yet in most organisations, the mapping from a lab extract to the LB domain is written fresh per study, because it exists as SAS or R code with study-specific paths baked into it, not as a declared relationship between a source column and a target variable with a transformation.

Certara Pinnacle 21 is the de facto validation standard in this space and deserves that position, but validation is diagnosis rather than treatment: it tells you what is non-conformant, it does not do the mapping. Pinnacle 21 Enterprise adds a metadata repository and governance, and it is a platform you configure to its model rather than a transformation engine you own. Certara Formedix is strong on metadata-driven study design and automation, with a similar platform commitment. The SAS Clinical Standards Toolkit gives a serious framework if you are a deep SAS shop with the internal expertise to maintain it, which is a real if. Instem brings tooling with a services heritage. What none of them hands you is your own portfolio-wide mapping library that becomes more valuable every time you file.

What a custom build does: it separates the declaration from the execution. A mapping is metadata, source dataset and column, target domain and variable, transformation rule, controlled terminology assignment, origin and derivation text for the Define. The engine executes it and produces both the dataset and the documentation from the same source of truth, which is the property that removes the six-week loop. When study eight arrives with a familiar lab layout, you reuse a mapping set rather than a program, and the difference is measured in days.

Problem 2: Define-XML written by hand is always slightly wrong

Define-XML is the reviewer's map of your data, and version 2.1 with its accompanying reviewer guides is what agencies expect alongside standardised study data described in the FDA Study Data Technical Conformance Guide. When the Define is produced by a separate manual process, it drifts from the datasets. A variable gets a length change, a codelist gains a term, an origin changes from CRF to derived, and the Define says otherwise. Reviewers notice, and each question costs calendar time you do not have near a filing.

Generate it. If the mapping metadata is authoritative, the Define is a rendering of that metadata and cannot disagree with the data it describes. The same is true of the study data reviewer guide and the analysis data reviewer guide: the structural sections come from metadata, and the narrative sections are drafted from that structure for a human programmer to edit and sign. That is a fair use of language generation, writing from recorded facts rather than inventing them, and it typically removes days of assembly work per submission.

Problem 3: controlled terminology moves, and sponsor extensions are ungoverned

CDISC controlled terminology is published on a regular cycle, and your studies are pinned to different versions. On top of that sit your own extensible codelists, and in most organisations those extensions accumulate without governance, so the same concept appears three ways across three studies and nobody notices until an integrated summary is attempted and the pooling fails.

Treat terminology as versioned reference data with a promotion process. Published packages load automatically. Sponsor extensions require a named owner, a definition and an approval, and the system reports where an extension is used and whether a published term has since arrived that should replace it. The payoff is not conformance, it is that integration across studies stops being an archaeology project, which matters most exactly when you are assembling a summary of safety across a programme.

Problem 4: ADaM traceability is a claim most organisations cannot substantiate

ADaM's core promise is traceability from analysis results back through analysis datasets to SDTM to the CRF. In practice that trail is asserted in a document and implemented in code that only its author fully understands. When a reviewer asks how a particular analysis flag was derived for a specific subject, the answer requires someone to read a program.

Build derivation as declared logic with the derivation text and the executable rule bound together, so the ADaM specification and the code cannot diverge. Then a subject-level trace becomes a query: this flag came from these SDTM records under this rule at this version. Double programming remains a good practice and should be supported rather than replaced, with the system comparing independent outputs and reporting differences rather than a human eyeballing listings.

Problem 5: legacy and acquired studies arrive in structures nobody planned for

An acquisition brings six studies. A partner delivers data in their own convention. An old study predates your standards entirely. Each one is a discovery exercise before it is a mapping exercise, and that discovery is where the schedule risk sits.

This is where a model earns its place, and the honest framing matters. Feed it the source column names, sample values, any available annotated CRF and your existing mapping library, and it proposes candidate mappings ranked by confidence. A programmer accepts, rejects or edits every one. The value is not accuracy on any single mapping, it is that a programmer starts from a populated draft rather than a blank specification, and that the draft is informed by every mapping your organisation has already approved. In our experience this is where the calendar time actually comes back on legacy work. What the model must never do is derive data. Derivations run as declared, versioned, testable rules, because a regulator is entitled to see exactly how a number was produced and no probabilistic answer is acceptable there.

Problem 6: the standard version changes and you have no regeneration path

Agency data standards catalogues move. A new SDTMIG version becomes expected for studies starting after a date. If mappings are programs, that migration is a manual pass over every study. If mappings are metadata with the target standard as a parameter, regeneration is a run followed by a review of what changed. That single property is the strongest long-term argument for owning this layer, because standards will keep moving for as long as you keep filing.

What this costs and how long it takes

A first release with a metadata repository, mapping declaration and execution engine, generated Define-XML, controlled terminology versioning and an automated conformance loop runs $85,000 to $170,000 and ships in 12 to 18 weeks, based on Digital Heroes delivery experience. A full platform adding ADaM derivation traceability, double programming comparison, reviewer guide generation, legacy onboarding with model-assisted mapping proposals, sponsor terminology governance and regeneration across standard versions runs $220,000 to $550,000 phased over 8 to 14 months.

What drives price up specifically here: the number of source systems and CRO delivery conventions you must accommodate, since each convention is its own onboarding pattern. Whether you run SAS, R or both, because a shop maintaining two execution paths is doing double the validation work. Integration with the CDISC Library for standards metadata rather than maintaining your own copies. Therapeutic area user guides, if you work in areas with specific supplemental standards. And validation, because a system producing submission datasets is a regulated computerised system under 21 CFR Part 11 and needs qualification. What keeps it down: implementing three representative studies end to end before widening, and resisting the temptation to model every domain you have ever seen in release one.

Build versus buy, and when licensing is clearly correct

Licence, and do not call us, if you file rarely and your studies are conventional. Pinnacle 21 Enterprise plus experienced contract programmers will get you a compliant package for less than any build, and there is no portfolio effect to capture because you do not have a portfolio. Licence also if your organisation has no standards governance function, because tooling does not create governance and an ungoverned metadata repository decays faster than a folder of programs.

Build when two or more of these are true. You file more than a couple of submissions a year and the same mapping work keeps repeating. You have acquired assets in unfamiliar structures and legacy onboarding has become a recurring cost line. You receive data from multiple CROs on different conventions. Your controlled terminology extensions are ungoverned and integration across studies is painful. Or your Define-XML is assembled by hand near the filing date, which is a schedule risk hiding as a task. The tipping point is that mapping knowledge is intellectual property, and the organisations that capture it as metadata file faster every year while the ones that keep it as programs file at the same speed forever.

How to choose a developer for CDISC standards tooling

Ask them to explain how the Define-XML stays synchronised with the datasets. If the answer involves a separate generation step run late in the process, they are rebuilding the problem you are trying to leave. The correct answer is that both are rendered from the same metadata and cannot diverge.

Ask what happens when SDTMIG moves a version. A partner who has done this talks about the target standard as a parameter and about regeneration with a change report. A partner who talks about updating the programs has not thought past the first submission.

Ask how derivations are tested. Fixtures with expected outputs, run automatically, is the only answer that holds up when a reviewer asks how a flag was derived. Then confirm in writing before kickoff that you own the repository, the metadata, the infrastructure accounts and the validation package. Your mapping library will become one of the more valuable assets in the department, and it must be portable to any partner including away from us. At Digital Heroes it belongs to the client from the first commit.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. McKinsey found that tech debt can amount to 20-40% of the value of a company's entire technology estate before depreciation, and CIOs report that 10-20% of the budget for new products is diverted to resolving tech-debt issues. Source: McKinsey & Company (2020) →
  2. Only 16% of respondents said their organizations' digital transformations had successfully improved performance and equipped them to sustain gains over the long term; even in digitally savvy industries such as high tech, media, and telecom, self-reported success rates did not exceed 26%. Source: McKinsey & Company (2018) →
  3. Brandon Hall Group research on onboarding reports that done well, structured onboarding drives measurable gains in new-hire productivity, employee engagement, and retention; the page notes 41% of organizations experience greater than 5% turnover among new hires. Source: Brandon Hall Group (2024) →
  4. Total US training expenditure rose 4.9% to $102.8 billion; learning management systems were used at 89% of organizations (90% of large, 97% of midsize, 84% of small companies), with average training at 40 hours per employee and $874 spent per learner. Source: Training Magazine (2025) →
Ishaan C. · Shopify Plus Tech Lead · Delhi

Ishaan is the technical lead on Shopify Plus builds at Digital Heroes, working on checkout extensions, custom apps, integrations with ERP and the parts of a store that outgrow standard themes. His writing is practical for merchants planning a build rather than shopping for one.

View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.

FAQ

Frequently asked questions

How much does it cost to build custom SDTM and ADaM mapping software?
A first release with a metadata repository, a mapping declaration and execution engine, generated Define-XML, controlled terminology versioning and an automated conformance loop runs $85,000 to $170,000 and ships in 12 to 18 weeks, based on Digital Heroes delivery experience. A full platform adding ADaM derivation traceability, double programming comparison, reviewer guide generation and legacy study onboarding runs $220,000 to $550,000 phased over 8 to 14 months. Supporting both SAS and R execution paths is the most common hidden cost.
Does Pinnacle 21 already solve this, or do we still need to build?
Pinnacle 21 is the de facto validation standard and it is worth having, but validation is diagnosis rather than treatment: it reports non-conformance, it does not perform your mapping. Pinnacle 21 Enterprise adds a metadata repository and governance you configure to its model. What neither gives you is a portfolio-wide mapping library that gets more valuable each time you file, which is the actual asset if you submit regularly or keep acquiring studies in unfamiliar structures.
How do you stop Define-XML from drifting away from the datasets?
Generate both from the same metadata. If the mapping declaration is authoritative, holding source, target, transformation, controlled terminology assignment, origin and derivation text, then the datasets and the Define-XML are two renderings of one truth and cannot disagree. Hand-assembled Define files drift the moment a length changes or a codelist gains a term, and each reviewer question about that drift costs calendar time exactly when you have none.
Can AI generate SDTM mappings automatically?
It can propose them, which is genuinely useful, and it must not derive data. Given source column names, sample values, any annotated CRF and your existing approved mapping library, a model produces ranked candidate mappings so a programmer starts from a populated draft rather than a blank specification. Every proposal is accepted, rejected or edited by a human. Derivations themselves run as declared, versioned, testable rules, because a regulator is entitled to see exactly how a number was produced.
How should sponsor controlled terminology extensions be governed?
As versioned reference data with a promotion process, not as ad hoc additions inside study programs. Published CDISC packages should load automatically, and each sponsor extension should require a named owner, a definition and an approval, with reporting on where it is used and whether a published term has since arrived that supersedes it. Ungoverned extensions are invisible until you attempt an integrated summary across a programme, at which point pooling fails and the cleanup is archaeology.
What happens to our mappings when a new SDTMIG version becomes expected?
If mappings live as programs, migration is a manual pass over every study and it will consume a quarter. If mappings live as metadata with the target standard as a parameter, migration is a regeneration run followed by review of a change report. Agency data standards catalogues keep moving, so this property compounds: it is the strongest long-term argument for owning the mapping layer rather than rebuilding it study by study.
How do you prove ADaM traceability to a reviewer?
By binding the derivation text and the executable rule together so the specification and the code cannot diverge, then exposing a subject-level trace as a query rather than as a program-reading exercise. A reviewer asking how a particular analysis flag was set for one subject should get an answer showing the contributing SDTM records, the rule applied and the rule version. Double programming remains good practice, with the system comparing independent outputs and reporting differences automatically.
Does standards tooling need 21 CFR Part 11 validation?
Yes if it produces the datasets and documentation you submit. Treat it as a regulated computerised system with a validation plan, requirements traced to executed test scripts, qualification, documented change control and periodic review, using GAMP 5 as the framework. The practical benefit of a metadata-driven design is that revalidation after a change is narrower, because you are re-testing a declared rule and its fixtures rather than a program someone rewrote by hand.
We acquired six studies with unfamiliar data structures. Where do we start?
Start with discovery rather than mapping. Profile every source dataset, catalogue column names, value distributions and any annotated CRF, then let a model propose candidate mappings against your existing approved library for a programmer to confirm. Onboard the studies most likely to appear in a filing or an integrated summary first, and accept that some legacy assets are best left as they are with a documented rationale. The schedule risk in acquired studies is discovery, not transformation.
How much does a custom internal tool cost to build?
Most custom internal tools cost $8,000 to $40,000 to build, based on Digital Heroes delivery data across 2,000+ client projects. A single-purpose tool like an approval dashboard or inventory tracker sits at the low end, while a multi-department platform with role-based access and several integrations pushes past $40,000. The three biggest cost drivers are the number of user roles, the number of systems the tool must connect to, and custom reporting requirements.
At what point does Retool cost more than building a custom tool?
The crossover usually lands between 25 and 50 daily users. At Retool's published Business rates of $50 per standard user and $15 per end user monthly, a 40-person deployment with a typical seat mix runs roughly $9,000 to $15,000 per year, every year, while a comparable custom tool built once for $20,000 to $30,000 carries no per-seat fees and costs about 15 to 20 percent of the build price annually to maintain. On a three-year horizon, custom comes out ahead for most growing teams in Digital Heroes engagements.
How do I know when spreadsheets are no longer enough to run my operations?
Replace the spreadsheet once more than three people edit it, versions travel by email, or a single broken formula could cost real money. Other reliable signals: staff keep personal shadow copies, month-end reporting takes days of manual assembly, and nobody can say who changed a number or why. In Digital Heroes discovery calls the tipping point is almost always a specific expensive error, a mispriced quote, a missed order, or payroll built on a tab someone sorted wrong.
Who owns the code when an agency builds my software?
You should, completely, through a written intellectual property assignment that transfers everything on final payment; without that clause, copyright stays with whoever wrote the code by default. Insist that the repository lives in your own GitHub organization from day one and that hosting, domains, and third-party accounts are registered to you. Also check for licenses to the agency's proprietary frameworks buried in the contract, because those can make switching vendors practically impossible even when you own your own code.
Is a custom internal tool secure enough for HR records and financial data?
A properly built custom tool is generally safer for sensitive data than the shared spreadsheet it replaces, because you get role-based access, audit logs, encrypted storage, and the ability to cut one person's access instantly. Ask the agency specifically for encryption in transit and at rest, permissions down to the field level, and an audit trail showing who viewed or changed each record. If HIPAA, GDPR, or SOC 2 expectations from enterprise clients apply to you, raise it before the quote, because compliance features add real scope.
What should I prepare before contacting a software development agency?
A one-page brief beats a 40-page requirements document: the business problem in plain words, who will use the system, the 5 to 10 workflows it must handle, the tools it must connect to, and your budget range and deadline driver. You do not need wireframes, a specification, or technical vocabulary; producing those is the agency's job during discovery. Stating a budget range up front is the single best move, because it gets you honest scoping instead of a quote engineered to win the meeting.
How long does it take to build an internal tool from scratch?
A working first version typically ships in 4 to 8 weeks, and larger multi-module tools run 10 to 16 weeks. Across Digital Heroes internal tool projects the schedule splits into roughly one week of process mapping, 3 to 6 weeks of build, and 1 to 2 weeks of testing with your actual staff. The most common delay is not development but waiting on the client for sample data and workflow decisions, so name one internal owner before kickoff.
What does it cost to keep an internal tool running after launch, and do we need to hire a developer?
Budget 15 to 20 percent of the build cost per year, so a $25,000 tool runs roughly $300 to $400 a month covering hosting, security patches, dependency updates, and small tweaks, figures drawn from Digital Heroes maintenance contracts. You do not need an in-house developer; a monthly retainer with the agency that built it covers the typical internal tool comfortably. Hosting itself is cheap for internal audiences, often $20 to $100 a month, because you serve dozens of users rather than the open internet.
Will a custom internal tool scale as our company grows?
Yes, provided it sits on a standard stack with a real database: PostgreSQL comfortably handles millions of records, and adding users costs hosting pennies rather than per-seat fees. The real scaling risks are organizational, not technical: new departments want features, processes change, and the tool needs a budget line to evolve. Set aside a small quarterly improvement budget instead of treating launch as the finish line, and the tool stays useful for a decade rather than getting rebuilt every two years.
Who owns the code when an agency builds our internal tool?
You should, outright, with full IP transfer in the contract and the code delivered to a repository you control, such as your own GitHub organization. Digital Heroes transfers complete ownership on final payment as standard practice, and any agency that keeps the code or licenses it back to you is building a dependency you will pay for later. Confirm you also own the hosting, domain, and database accounts, since many of the vendor disputes Digital Heroes gets called into involve infrastructure registered under the agency's name.
How do I calculate the ROI of a custom internal tool?
Count hours first: multiply the weekly hours staff spend on the manual process by their loaded hourly cost, then add the cost of errors such as mispriced quotes or missed renewals. A tool saving a 10-person team 5 hours each per week recovers about 2,500 hours a year, which repays a $20,000 to $30,000 build well inside a year at typical wages. Most internal tools Digital Heroes delivers reach payback in 6 to 18 months, with quoting and billing tools at the fast end because they plug revenue leaks, not just time.
Who can build a custom internal tools system?

Digital Heroes builds custom internal tools systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other internal tools companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading
let's build

Build something worth launching.

A plan, a team, a timeline, within 24 hours. No decks, no discovery calls. Tell us what you're building and we'll come back with a real scope and a real number.

message us directly · we reply within one business day

mission briefing

Monthly dispatch

Playbooks, real build costs, and what we're shipping. One email a month. No fluff.

visit us

New York HQ

1140 Broadway, Suite 704 · New York, NY 10001

Get directions
Online now

Hey there 👋 How can we help you today?