Compound Registration and Assay Data Software: Why Your Structure Activity Relationships Cannot Be Trusted
If you have more than about 30 bench chemists and biologists, register several thousand new compounds a year, and your assay results reach the structure activity table through a chain of Excel files, a custom registration and screening data platform is usually justified. A first release covering structure normalisation, registration rules, batch and lot identity and plate result loading typically runs $120,000 to $250,000 and ships in 16 to 24 weeks in Digital Heroes delivery experience. A full discovery data platform adding inventory, biologics entities, curve fitting, project dashboards and notebook integration lands at $300,000 to $750,000 phased over 12 to 20 months. Below about 15 scientists on small molecules only, CDD Vault costs a fraction of that and is the right answer.
Why the structure activity table is the thing that actually breaks
A medicinal chemist looks at a table of analogues and decides what to make next. That decision is the entire output of a discovery organisation, and it rests on a chain that almost nobody inspects: the structure as drawn, the salt form and the batch that was actually submitted, the plate well that batch went into, the raw readings from the reader, the curve fit, and the number that lands in the table. Break any link and the chemist optimises toward noise.
The break is usually mundane. A batch was resubmitted at higher purity and the assay ran on the old material. A plate map was pasted one row off. Two chemists registered what they thought were different compounds and they were the same parent with different counterions. A biologist reported an inhibitory concentration from a four point curve that should have been flagged as unfit. None of these is a dramatic failure. All of them quietly move a programme in the wrong direction for months, and by the time someone re tests, the series has been abandoned or defended on bad evidence.
Registration software exists to make identity unambiguous. Screening data software exists to make the attachment of a result to that identity unambiguous. Everything else in a discovery informatics stack is convenience on top of those two jobs.
What Dotmatics, CDD Vault, Benchling and Revvity Signals actually leave to you
These are serious products with real chemistry underneath. Dotmatics has depth across registration, assay data and querying and is the default answer for a mid size chemistry organisation. CDD Vault is well built, genuinely inexpensive relative to the category, and correct for smaller teams. Benchling is strong on biologics entity modelling and has become the default in biotech. Revvity Signals and Schrodinger LiveDesign both bring capable analysis layers. If your science fits their entity models, buying beats building and we will say so.
Three gaps recur. The first is registration business rules. Every organisation has its own conventions: what counts as a new compound versus a new batch, how salts and solvates are handled, how mixtures and undefined stereochemistry are treated, what happens when a structure is corrected after results exist, who is allowed to override a duplicate warning. Packaged systems ship one opinionated rule set with configuration around the edges, and organisations with a twenty year legacy of their own conventions cannot adopt someone else's without invalidating historical identity.
The second is modality. Small molecule registration is a solved shape. Peptides with non natural residues, antibody drug conjugates with linker and payload and drug to antibody ratio, oligonucleotides with modified backbones, PROTACs with two binding elements, and cell lines or strains as first class entities all need models that do not exist in a small molecule product. Bolting them on as text fields is what most teams end up doing, and it kills every query that matters.
The third is the instrument and plate layer. Your specific readers, your specific plate formats, your handling of 384 and 1536 well layouts, your controls placement, your normalisation convention and your curve fitting rules are local. Vendors provide parsers for common formats and stop at the boundary of your screening cascade.
Registration rules are the part everyone underestimates
Structure normalisation sounds like a solved library call. It is not a policy. Before you can decide whether an incoming structure is new, you have to decide how you standardise it: tautomer handling, charge and salt stripping, stereochemistry perception, isotopes, whether a defined single enantiomer and its racemate are the same registration or different ones. Then you generate an identity key, and an international chemical identifier such as InChI is a reasonable basis, but the standardisation decisions in front of it determine everything.
The organisational half is harder. A chemist submits a structure and the system says it already exists. Sometimes that is correct and useful. Sometimes it is a false match caused by a normalisation rule the chemist disagrees with, and the workflow has to allow a documented override with a reviewer, because a system that cannot be overridden gets bypassed by a spreadsheet within a month. Corrections are the sharpest edge: a structure proven wrong after two years of assay data must be correctable without orphaning that data or silently changing the meaning of published results. That is a versioned identity problem, and it is the single clearest reason organisations build rather than buy.
Batch identity is where assay data goes wrong
Results attach to batches, not to compounds. That sentence is the whole discipline. A parent compound may have twelve batches with different purity, different salt forms, different suppliers and different solid state. If the assay result is attached to the parent, a bad batch contaminates the whole series, and nobody can find out why one number does not fit.
A build has to carry the parent to batch to sample to well chain end to end. Registration produces the batch. Inventory produces a sample, meaning a specific vial or a specific plate well at a specific concentration in a specific solvent. The plate map records which sample went into which well, and it must come from the liquid handler's actual output rather than a chemist's intended layout, because those two are not always the same. Then the reader file lands, wells are matched, controls are located, normalisation runs, and the curve is fitted with rules that decide when a fit is not reportable. When someone questions a number six months later, that entire chain should be visible on one screen. That single capability is what most spreadsheet based organisations are actually paying for.
What a custom build must include
- Structure standardisation with your normalisation policy made explicit and versioned, not hidden in a library default.
- Registration workflow with duplicate detection, documented override, reviewer roles and corrections handled as versioned identity.
- Parent, batch, sample and well as separate first class objects with a traceable chain.
- Modality specific entities for peptides, conjugates, oligonucleotides and biologics where your pipeline needs them, rather than text fields.
- Inventory with location, amount, concentration and depletion, because a chemist asking whether material exists is the most frequent query in the building.
- Plate map ingestion from liquid handler output, with control placement and layout templates per assay.
- Reader file parsers for your specific instruments, normalisation and curve fitting with reportability rules.
- Assay definitions with protocol versions, so a result knows which version of the assay produced it.
- Query and visualisation that a chemist uses without help, including substructure and similarity search that returns in seconds on your full collection.
- An audit trail, because invention records and data integrity questions both depend on knowing who changed what.
What it costs and how long it takes
Across the projects Digital Heroes has delivered in scientific data platforms, a first release covering registration, batch identity and plate result loading runs $120,000 to $250,000 and ships in 16 to 24 weeks. A full discovery platform adding inventory, biologics entities, curve fitting, dashboards and notebook integration runs $300,000 to $750,000 across 12 to 20 months.
What drives the number up specifically in discovery informatics: the number of modalities, because each new entity type is a data model and a user interface, not a field. Legacy migration, which is the largest hidden cost, since re registering twenty years of compounds under a new normalisation policy will surface thousands of conflicts that a human chemist has to adjudicate. Instrument parser count. Substructure search performance at collection scale, which needs a proper chemical cartridge and index rather than a filter over a table. And integration with an existing electronic lab notebook, inventory system or ordering platform.
What keeps it down: one modality, your top ten assays, and an honest decision to leave the historical collection in the old system read only rather than migrating everything on day one.
Build versus buy, stated plainly
Buy CDD Vault if you are under about 15 scientists on small molecules. It is inexpensive, it works, and a custom build at that size is a distraction from making compounds. Buy Dotmatics or Benchling if your entity model matches theirs reasonably well and you would rather spend your engineering attention elsewhere. Both are defensible decisions for organisations several times larger than that.
Build when two or more of these are true. Your registration conventions have twenty years of history and adopting a vendor's rules would break historical identity. You work across modalities that no single product models properly, which is now common in biotech. Your screening cascade has assay specific normalisation and reportability rules that live in one scientist's spreadsheet. You are paying per seat for a system that half your organisation cannot use, so they keep working in Excel. Or your query performance on the full collection is bad enough that chemists have stopped asking questions, which is the most expensive failure of all and the hardest to see.
How to choose a developer for chemistry data software
Ask them to explain the difference between a compound, a batch and a sample, and what happens when a structure is corrected after results exist. A developer who has done this work answers immediately and mentions versioning. A developer who talks about products and records has built a catalogue and is about to learn chemistry on your budget.
Ask what chemistry toolkit they intend to use and why. RDKit, OpenBabel and commercial toolkits have different behaviour on tautomers and stereochemistry perception, and the choice has consequences for your identity keys that last forever. A developer with no opinion here has not thought about it.
Ask how substructure search will be implemented and what performance they will commit to at your collection size. If the answer does not involve fingerprint screening and a proper index, expect queries that time out.
Ask who owns the code and get it in writing before kickoff. You should hold the repository, the infrastructure accounts and the right to hire anyone else to continue. At Digital Heroes the code is yours from the first commit. This matters more than usual here, because a registration system becomes the memory of the research organisation and you cannot afford to have that memory hosted by a vendor you have fallen out with.
The evidence behind this guide
Independent findings on why this investment pays off. Every link goes to the primary source.
- Only 22% of firms are 'future ready' having significantly transformed digitally; these companies show average revenue growth 17.3 percentage points and net margins 14.0 percentage points above their industry average. Source: MIT Center for Information Systems Research (MIT Sloan) (2022) →
- Senior executives report the highest average compensation among developer roles (e.g., $225K median in the US), and reported salary bands shifted downward year-over-year ($60-75K vs. $70-85K in 2023), underscoring how compensation varies sharply by role and location. Source: Stack Overflow (2024) →
- In the Flexera 2025 State of ITAM report, respondents reported roughly 33% of SaaS spend is wasted, underscoring how paying for off-the-shelf seats and tiers that go unused erodes the supposed cost advantage of generic SaaS. Source: Flexera (2025) →
- Gallup reports global employee engagement fell to 20% in 2025 (its lowest since 2020, down from a 2022-2023 peak of 23%), and estimates low engagement costs the world economy an estimated $10 trillion in lost productivity, or 9% of global GDP. (Note: this figure appears in Gallup's evergreen State of the Global Workplace page, currently reflecting the 2026 edition reporting on 2025 data.). Source: Gallup (2025) →
Eleanor leads client services across the UK and EU, which means she sits between what a client asks for and what the delivery teams can realistically build. She writes about scoping, budget conversations and the questions worth asking before a build starts.
View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.
Frequently asked questions
How much does a custom compound registration and screening system cost?
Is CDD Vault or Dotmatics enough, or should we build?
Why do assay results have to attach to a batch rather than a compound?
How should structure normalisation and duplicate detection work?
Can a custom system handle peptides, conjugates and other modalities?
How long does it take to build a discovery informatics platform?
Do we have to migrate our entire historical compound collection?
Why is substructure search performance a design decision rather than a feature?
Who owns the code if an agency builds our registration system?
How do we get years of data out of our old system and into the new one?
Our developer disappeared mid-project. Can another team pick up the code?
We run everything on Airtable and spreadsheets. When is it time to go custom?
What should I have ready before I contact a development agency?
What is the biggest mistake first-time software buyers make?
Does it matter which tech stack the agency wants to use?
Can I build my product on a no-code tool like Bubble instead of hiring developers?
How do I make sure custom software is secure and compliant with rules like HIPAA?
How long does it take from first call to software my team can actually use?
Will custom software work with the tools we already use, like QuickBooks and Stripe?
Who can build a custom software system?
Digital Heroes builds custom software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.
Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.
What makes Digital Heroes different from other software companies?
Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.
Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.
How can I check Digital Heroes is legitimate before getting in touch?
Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.
Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.