Industry guide · Custom Software

Compound Registration and Assay Data Software: Why Your Structure Activity Relationships Cannot Be Trusted

Compound Registration and Screening software visual showing atom, layout grid, and data records.
The short answer

If you have more than about 30 bench chemists and biologists, register several thousand new compounds a year, and your assay results reach the structure activity table through a chain of Excel files, a custom registration and screening data platform is usually justified. A first release covering structure normalisation, registration rules, batch and lot identity and plate result loading typically runs $120,000 to $250,000 and ships in 16 to 24 weeks in Digital Heroes delivery experience. A full discovery data platform adding inventory, biologics entities, curve fitting, project dashboards and notebook integration lands at $300,000 to $750,000 phased over 12 to 20 months. Below about 15 scientists on small molecules only, CDD Vault costs a fraction of that and is the right answer.

Why the structure activity table is the thing that actually breaks

A medicinal chemist looks at a table of analogues and decides what to make next. That decision is the entire output of a discovery organisation, and it rests on a chain that almost nobody inspects: the structure as drawn, the salt form and the batch that was actually submitted, the plate well that batch went into, the raw readings from the reader, the curve fit, and the number that lands in the table. Break any link and the chemist optimises toward noise.

The break is usually mundane. A batch was resubmitted at higher purity and the assay ran on the old material. A plate map was pasted one row off. Two chemists registered what they thought were different compounds and they were the same parent with different counterions. A biologist reported an inhibitory concentration from a four point curve that should have been flagged as unfit. None of these is a dramatic failure. All of them quietly move a programme in the wrong direction for months, and by the time someone re tests, the series has been abandoned or defended on bad evidence.

Registration software exists to make identity unambiguous. Screening data software exists to make the attachment of a result to that identity unambiguous. Everything else in a discovery informatics stack is convenience on top of those two jobs.

What Dotmatics, CDD Vault, Benchling and Revvity Signals actually leave to you

These are serious products with real chemistry underneath. Dotmatics has depth across registration, assay data and querying and is the default answer for a mid size chemistry organisation. CDD Vault is well built, genuinely inexpensive relative to the category, and correct for smaller teams. Benchling is strong on biologics entity modelling and has become the default in biotech. Revvity Signals and Schrodinger LiveDesign both bring capable analysis layers. If your science fits their entity models, buying beats building and we will say so.

Three gaps recur. The first is registration business rules. Every organisation has its own conventions: what counts as a new compound versus a new batch, how salts and solvates are handled, how mixtures and undefined stereochemistry are treated, what happens when a structure is corrected after results exist, who is allowed to override a duplicate warning. Packaged systems ship one opinionated rule set with configuration around the edges, and organisations with a twenty year legacy of their own conventions cannot adopt someone else's without invalidating historical identity.

The second is modality. Small molecule registration is a solved shape. Peptides with non natural residues, antibody drug conjugates with linker and payload and drug to antibody ratio, oligonucleotides with modified backbones, PROTACs with two binding elements, and cell lines or strains as first class entities all need models that do not exist in a small molecule product. Bolting them on as text fields is what most teams end up doing, and it kills every query that matters.

The third is the instrument and plate layer. Your specific readers, your specific plate formats, your handling of 384 and 1536 well layouts, your controls placement, your normalisation convention and your curve fitting rules are local. Vendors provide parsers for common formats and stop at the boundary of your screening cascade.

Registration rules are the part everyone underestimates

Structure normalisation sounds like a solved library call. It is not a policy. Before you can decide whether an incoming structure is new, you have to decide how you standardise it: tautomer handling, charge and salt stripping, stereochemistry perception, isotopes, whether a defined single enantiomer and its racemate are the same registration or different ones. Then you generate an identity key, and an international chemical identifier such as InChI is a reasonable basis, but the standardisation decisions in front of it determine everything.

The organisational half is harder. A chemist submits a structure and the system says it already exists. Sometimes that is correct and useful. Sometimes it is a false match caused by a normalisation rule the chemist disagrees with, and the workflow has to allow a documented override with a reviewer, because a system that cannot be overridden gets bypassed by a spreadsheet within a month. Corrections are the sharpest edge: a structure proven wrong after two years of assay data must be correctable without orphaning that data or silently changing the meaning of published results. That is a versioned identity problem, and it is the single clearest reason organisations build rather than buy.

Batch identity is where assay data goes wrong

Results attach to batches, not to compounds. That sentence is the whole discipline. A parent compound may have twelve batches with different purity, different salt forms, different suppliers and different solid state. If the assay result is attached to the parent, a bad batch contaminates the whole series, and nobody can find out why one number does not fit.

A build has to carry the parent to batch to sample to well chain end to end. Registration produces the batch. Inventory produces a sample, meaning a specific vial or a specific plate well at a specific concentration in a specific solvent. The plate map records which sample went into which well, and it must come from the liquid handler's actual output rather than a chemist's intended layout, because those two are not always the same. Then the reader file lands, wells are matched, controls are located, normalisation runs, and the curve is fitted with rules that decide when a fit is not reportable. When someone questions a number six months later, that entire chain should be visible on one screen. That single capability is what most spreadsheet based organisations are actually paying for.

What a custom build must include

  • Structure standardisation with your normalisation policy made explicit and versioned, not hidden in a library default.
  • Registration workflow with duplicate detection, documented override, reviewer roles and corrections handled as versioned identity.
  • Parent, batch, sample and well as separate first class objects with a traceable chain.
  • Modality specific entities for peptides, conjugates, oligonucleotides and biologics where your pipeline needs them, rather than text fields.
  • Inventory with location, amount, concentration and depletion, because a chemist asking whether material exists is the most frequent query in the building.
  • Plate map ingestion from liquid handler output, with control placement and layout templates per assay.
  • Reader file parsers for your specific instruments, normalisation and curve fitting with reportability rules.
  • Assay definitions with protocol versions, so a result knows which version of the assay produced it.
  • Query and visualisation that a chemist uses without help, including substructure and similarity search that returns in seconds on your full collection.
  • An audit trail, because invention records and data integrity questions both depend on knowing who changed what.

What it costs and how long it takes

Across the projects Digital Heroes has delivered in scientific data platforms, a first release covering registration, batch identity and plate result loading runs $120,000 to $250,000 and ships in 16 to 24 weeks. A full discovery platform adding inventory, biologics entities, curve fitting, dashboards and notebook integration runs $300,000 to $750,000 across 12 to 20 months.

What drives the number up specifically in discovery informatics: the number of modalities, because each new entity type is a data model and a user interface, not a field. Legacy migration, which is the largest hidden cost, since re registering twenty years of compounds under a new normalisation policy will surface thousands of conflicts that a human chemist has to adjudicate. Instrument parser count. Substructure search performance at collection scale, which needs a proper chemical cartridge and index rather than a filter over a table. And integration with an existing electronic lab notebook, inventory system or ordering platform.

What keeps it down: one modality, your top ten assays, and an honest decision to leave the historical collection in the old system read only rather than migrating everything on day one.

Build versus buy, stated plainly

Buy CDD Vault if you are under about 15 scientists on small molecules. It is inexpensive, it works, and a custom build at that size is a distraction from making compounds. Buy Dotmatics or Benchling if your entity model matches theirs reasonably well and you would rather spend your engineering attention elsewhere. Both are defensible decisions for organisations several times larger than that.

Build when two or more of these are true. Your registration conventions have twenty years of history and adopting a vendor's rules would break historical identity. You work across modalities that no single product models properly, which is now common in biotech. Your screening cascade has assay specific normalisation and reportability rules that live in one scientist's spreadsheet. You are paying per seat for a system that half your organisation cannot use, so they keep working in Excel. Or your query performance on the full collection is bad enough that chemists have stopped asking questions, which is the most expensive failure of all and the hardest to see.

How to choose a developer for chemistry data software

Ask them to explain the difference between a compound, a batch and a sample, and what happens when a structure is corrected after results exist. A developer who has done this work answers immediately and mentions versioning. A developer who talks about products and records has built a catalogue and is about to learn chemistry on your budget.

Ask what chemistry toolkit they intend to use and why. RDKit, OpenBabel and commercial toolkits have different behaviour on tautomers and stereochemistry perception, and the choice has consequences for your identity keys that last forever. A developer with no opinion here has not thought about it.

Ask how substructure search will be implemented and what performance they will commit to at your collection size. If the answer does not involve fingerprint screening and a proper index, expect queries that time out.

Ask who owns the code and get it in writing before kickoff. You should hold the repository, the infrastructure accounts and the right to hire anyone else to continue. At Digital Heroes the code is yours from the first commit. This matters more than usual here, because a registration system becomes the memory of the research organisation and you cannot afford to have that memory hosted by a vendor you have fallen out with.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. Only 22% of firms are 'future ready' having significantly transformed digitally; these companies show average revenue growth 17.3 percentage points and net margins 14.0 percentage points above their industry average. Source: MIT Center for Information Systems Research (MIT Sloan) (2022) →
  2. Senior executives report the highest average compensation among developer roles (e.g., $225K median in the US), and reported salary bands shifted downward year-over-year ($60-75K vs. $70-85K in 2023), underscoring how compensation varies sharply by role and location. Source: Stack Overflow (2024) →
  3. In the Flexera 2025 State of ITAM report, respondents reported roughly 33% of SaaS spend is wasted, underscoring how paying for off-the-shelf seats and tiers that go unused erodes the supposed cost advantage of generic SaaS. Source: Flexera (2025) →
  4. Gallup reports global employee engagement fell to 20% in 2025 (its lowest since 2020, down from a 2022-2023 peak of 23%), and estimates low engagement costs the world economy an estimated $10 trillion in lost productivity, or 9% of global GDP. (Note: this figure appears in Gallup's evergreen State of the Global Workplace page, currently reflecting the 2026 edition reporting on 2025 data.). Source: Gallup (2025) →
Eleanor W. · VP Client Services · UK & EU · London

Eleanor leads client services across the UK and EU, which means she sits between what a client asks for and what the delivery teams can realistically build. She writes about scoping, budget conversations and the questions worth asking before a build starts.

View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.

FAQ

Frequently asked questions

How much does a custom compound registration and screening system cost?
A first release covering structure normalisation, registration rules, batch identity and plate result loading typically runs $120,000 to $250,000 and ships in 16 to 24 weeks, based on Digital Heroes delivery experience. A full discovery platform with inventory, biologics entities, curve fitting and dashboards runs $300,000 to $750,000 over 12 to 20 months. Legacy re registration of an existing collection is usually the largest hidden cost.
Is CDD Vault or Dotmatics enough, or should we build?
CDD Vault is the right answer for smaller small molecule teams, roughly under 15 scientists, and it is inexpensive for what it does. Dotmatics suits mid size chemistry organisations whose entity model matches theirs. Building becomes justified when your registration conventions carry decades of history that a vendor's rules would invalidate, when you work across modalities such as conjugates and oligonucleotides that no single product models properly, or when per seat licensing pushes half your team back into Excel.
Why do assay results have to attach to a batch rather than a compound?
Because different batches of the same parent can differ in purity, salt form, supplier and solid state, and attaching a result to the parent lets one bad batch contaminate the entire series. Keeping parent, batch, sample and plate well as separate linked objects means a surprising number can be traced back to the specific material and the specific well that produced it. Without that chain, nobody can explain why one analogue does not fit the trend.
How should structure normalisation and duplicate detection work?
Normalisation is a policy decision, not a library default: you have to decide how tautomers, charges, salts, stereochemistry and isotopes are treated before you generate an identity key such as an international chemical identifier. Duplicate detection then has to allow a documented override with a reviewer, because a system that cannot be overridden is bypassed with a spreadsheet within weeks. Write the policy down and version it, since changing it later reinterprets your whole collection.
Can a custom system handle peptides, conjugates and other modalities?
Yes, and this is one of the strongest arguments for building. Antibody drug conjugates need linker, payload and drug to antibody ratio as structured attributes, peptides need non natural residues, and oligonucleotides need modified backbones. Products built around small molecules tend to absorb these as text fields, which makes every downstream query useless. Each additional modality is a data model and user interface effort, so scope them deliberately rather than all at once.
How long does it take to build a discovery informatics platform?
A first release ships in 16 to 24 weeks in our experience. The critical path is rarely code. It is agreeing the registration policy across chemists who have held different conventions for years, and adjudicating the conflicts that surface when a legacy collection is re registered under one rule set. Organisations that assign a single decision maker for registration policy move dramatically faster than those that try to reach consensus.
Do we have to migrate our entire historical compound collection?
No, and forcing it on day one is a common way to stall a project. A practical pattern is to migrate the active collection and recent assay data, keep the legacy system available read only, and re register historical compounds in batches as programmes need them. Every migration wave surfaces normalisation conflicts that a chemist has to resolve by hand, so treating that adjudication as scheduled work rather than a surprise keeps the launch date honest.
Why is substructure search performance a design decision rather than a feature?
Because a naive implementation scans every structure and becomes unusable well before a collection reaches meaningful size, and chemists respond by no longer asking questions. Acceptable performance needs fingerprint based screening with a proper chemical index in front of the exact match step. Ask any developer to commit to a response time at your actual collection size before you sign, since retrofitting this later means changing the storage layer.
Who owns the code if an agency builds our registration system?
You should own the repository, the infrastructure accounts and the unrestricted right to hire another firm to continue the work, written into the contract before kickoff. At Digital Heroes the client owns the code from the first commit. This matters particularly for registration, because the system becomes the long term memory of your research organisation and you cannot have that memory locked inside a vendor relationship.
How do we get years of data out of our old system and into the new one?
Treat migration as a planned sub-project: a field-mapping document, at least one dry run on a copy of your data, then a cutover with the old system kept read-only for 30 days as a safety net. On Digital Heroes projects it consumes 10 to 15% of the budget when the old system has an export, and more when data must be pulled out screen by screen. Ask any vendor to walk you through their last migration before you sign.
Our developer disappeared mid-project. Can another team pick up the code?
Yes, this is a routine engagement, provided the code exists somewhere you can access, so your first move is securing the repository, hosting, and domain credentials today. A takeover starts with a one to two week paid code audit that ends in one of three verdicts: continue the build, keep the design but rebuild the weak parts, or start over. Digital Heroes has inherited enough projects to say plainly that sometimes the rebuild is cheaper than the rescue, and an honest agency will tell you which one you have before taking your money.
We run everything on Airtable and spreadsheets. When is it time to go custom?
The switch usually makes sense when you hit one of two walls: Airtable's record caps (125,000 records per base on the Business plan) or logic the tool cannot express, like multi-step approvals with conditional pricing. There is also a simple cost signal: 25 people on Business at roughly $45 per seat per month is about $13,500 a year, forever, for a tool you are already fighting. Custom is worth it when the workflow is core to how you make money; for peripheral processes, staying on Airtable is the right call.
What should I have ready before I contact a development agency?
Three things, none of them technical: a one-page description of the problem in your own words, a list of the tools and spreadsheets the new system must replace or connect to, and a must-have versus nice-to-have split of features. Add a budget range, even a wide one, because it changes the conversation from fantasy to engineering. You do not need a formal specification; producing that is what a discovery phase is for.
What is the biggest mistake first-time software buyers make?
Choosing the lowest quote without asking why it is the lowest. A bid 40% under the field usually gets there by skipping tests, documentation, and code review, which are invisible in a demo and brutal to pay for later; every stalled project Digital Heroes has been asked to rescue tells some version of that story. The second mistake is signing without a written scope, which reliably turns the winning cheap quote into 1.5x to 2x the price by launch.
Does it matter which tech stack the agency wants to use?
Yes, but not in the way most buyers expect: the goal is boring, popular technology such as React, Node.js or Python, and PostgreSQL, because any future team can maintain it and hiring a replacement developer takes days, not months. The red flag is an agency-proprietary framework or an unusual language, which welds you to that one vendor no matter what your contract says about code ownership. A useful test: could you find three freelancers fluent in this stack within a week? If not, push back.
Can I build my product on a no-code tool like Bubble instead of hiring developers?
For testing whether anyone wants the product, yes, and Bubble's paid plans start at $29 a month, which is the cheapest validation you will ever buy. The ceiling arrives with complex data relationships, heavy integrations, performance at a few thousand users, and the fact that you cannot export a Bubble app to servers you control. A path many Digital Heroes clients take: prove demand on no-code, then rebuild custom once revenue justifies it, treating the no-code version as a paid prototype rather than a foundation.
How do I make sure custom software is secure and compliant with rules like HIPAA?
Start with the baseline every business system should have: encryption in transit and at rest, role-based access control, and audit logs. If HIPAA applies, the hosting provider must sign a Business Associate Agreement, which AWS, Azure, and Google Cloud all offer, and access controls have to be designed in from day one, not bolted on. SOC 2 certifies a company's operating practices, not a codebase, so ask vendors what they have shipped in your regulated domain rather than which logos are on their website.
How long does it take from first call to software my team can actually use?
Plan for four to six months: two to three weeks of discovery, two to four weeks of design, then a 10 to 16 week build with testing. In Digital Heroes delivery experience the schedule killer is not engineering speed but decision lag; a client who takes two weeks to approve wireframes adds two weeks to launch. Book a weekly 30-minute decision slot before kickoff and most of that risk disappears.
Will custom software work with the tools we already use, like QuickBooks and Stripe?
Yes, and this is one of custom software's genuine advantages: QuickBooks, Stripe, Shopify, and most mainstream business tools publish documented APIs built for exactly this. Expect each standard integration to add one to two weeks of build time, and be suspicious of any quote that lists five integrations without asking what data flows in which direction. The hard cases are legacy systems with no API, which is a question to raise in discovery, not in week nine.
Who can build a custom software system?

Digital Heroes builds custom software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other software companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading
let's build

Build something worth launching.

A plan, a team, a timeline, within 24 hours. No decks, no discovery calls. Tell us what you're building and we'll come back with a real scope and a real number.

message us directly · we reply within one business day

mission briefing

Monthly dispatch

Playbooks, real build costs, and what we're shipping. One email a month. No fluff.

visit us

New York HQ

1140 Broadway, Suite 704 · New York, NY 10001

Get directions
Online now

Hey there 👋 How can we help you today?