Problems & solutions · Custom Software

Safety Data Sheet Authoring Software Problems: The 6 That Cost Real Money, and How to Avoid Them

SDS Authoring Chemical Compliance Software software overview illustration showing common problems and fixes.
The short answer

The most expensive failure is the reformulation that never triggers a document revision. A production chemist substitutes a surfactant because the original supplier is on allocation, the change goes through as a routine bill of materials revision, and the safety data sheet still describes a formulation you stopped making months ago. What that costs is not a fine in the first instance, it is a held shipment when the transport paperwork does not match the declared composition, a document recall across every distributor portal you have ever uploaded to, and a conversation in which you explain to a customer's health and safety team why your own record of your own product was wrong. That conversation costs accounts.

Why does building your own substance database go wrong so often?

The biggest scope failure in chemical compliance software is a developer proposing to build the regulatory content. It is an understandable mistake, because from the outside a substance list looks like a table. From the inside it is decades of curation: substance identifiers, harmonised classifications across jurisdictions, exposure limits, transport data, and the maintenance to keep all of it current as lists are updated.

Sphera, Verisk 3E and Chemwatch have spent a very long time assembling that, and a software budget spent recreating it buys you a worse version with no maintenance behind it. The failure usually surfaces about six months in, when the first list update lands and nobody has an answer for who applies it.

What is genuinely worth building is everything around the content: the connection to your live recipes, the approval workflow that reflects who in your organisation may sign off a hazard determination, the handling of customer specific variants, the phrase governance, and the generation of the right document for the right jurisdiction and language. Those are specific to how your business operates and no vendor can supply them.

So write the boundary into the specification before anyone quotes. Licensed content in, workflow and integration built, and an explicit answer to what happens when the licensor updates a classification. A proposal that does not name that boundary has not thought about it, and you will pay for the thinking later.

What goes wrong when you migrate legacy sheets and formulation data?

Two migrations, two very different difficulties. The document archive is the easy one to underestimate and the easy one to overdo. You have thousands of issued sheets across products, languages and versions, and the instinct is to import everything. Resist it. What you must retain is the issued record, meaning the exact document as supplied, with its issue date, so a superseded version can be produced if questioned. That is an archive, not a data set, and trying to parse old documents back into structured classification data produces confident errors.

The formulation data is the migration that actually determines your schedule. If recipes sit in the enterprise system as structured components with concentrations, integration is straightforward. In most manufacturers a meaningful share do not: some are text descriptions, some live in laboratory spreadsheets, some carry a trade name for a raw material rather than a composition, and a few reference a supplier blend whose composition you have never held.

Audit that before the project starts, not in week six. Count how many finished products have a fully structured composition, how many reference a raw material with no classification on file, and how many are described in prose. That count is the real project plan. Where composition is genuinely unavailable, decide explicitly whether you request it from the supplier, treat the material conservatively, or discontinue the product, and record which choice was made for each one.

Why do the enterprise system and supplier data integrations break after launch?

The single most common architectural error here is a nightly extract. A weekly or nightly file of formulations is not the same as knowing that a formulation changed at twenty past two on a Tuesday, and a product that changed and shipped the same day slips through a batch process entirely. Ask any prospective developer what happens in exactly that case. If the answer is that the change is picked up overnight, you have bought the original problem with better reporting.

The right shape is event driven from the system that holds the authoritative recipe, so a bill of materials revision immediately marks every affected sheet and label as stale and puts it in a work queue. That queue is the control. It is also the metric worth watching, because a queue that is growing tells you something about your change rate that no compliance report will.

The inbound side breaks differently. Supplier safety data sheets arrive as documents, by email, in whatever format the supplier chose, and they change without notice. Any process that depends on someone reading them will fail silently. Build an inbound register with an owner per supplier, a recorded date of the version you hold, and a cascade so a revised component classification recomputes every finished product containing it.

Then alert on absence as well as error. A feed that stops producing is more dangerous than one that fails loudly, because nothing looks wrong.

What happens when distribution and revision notification are not covered?

Issuing a corrected sheet is not compliance. Getting it to everyone who holds the previous one is compliance, and most manufacturers cannot say who that is. Sheets go out attached to order confirmations, get uploaded to customer portals by sales staff, and get emailed on request, with no register anywhere. So a revision becomes a broadcast to a mailing list that has been decaying for years, and the customer who most needs the update is the one whose contact left in the meantime.

The consequence is not theoretical. When a customer's health and safety team finds they are working to a superseded document, the question that follows is not about the document, it is about whether your quality system works at all. That is an audit question and it tends to arrive with a supplier questionnaire attached.

Covering it means treating distribution as a tracked event rather than an attachment. Every issue of a sheet to a party is recorded with the version, the date and the channel, whether that party is a customer, a distributor, a portal or one of your own sites. A revision then produces a precise notification list and evidence that notification happened, which is what you need when someone asks.

The same structured classification should generate any required notification of mixture information to a poison centre body, including the unique formula identifier the European scheme requires, rather than having that assembled separately by hand into a different format.

Should you build custom or configure what you already own?

If you make a few hundred formulations shipping into two or three jurisdictions with a stable recipe set, do not build. Chemwatch will produce compliant documents for far less than a build costs, and outsourcing authoring entirely to a service such as Verisk 3E is a legitimate answer at that scale. Spend the money on your laboratory instead, because the constraint on that business is formulation capability, not document generation.

If you already run the environment, health and safety module inside your enterprise system, look hard at what it can do before you build alongside it. It has the structural advantage of living where the recipes already are, which is the exact connection everyone else has to engineer. The honest test is whether your approval workflow and your customer specific variants fit inside it, or whether they are currently being handled by email around the product.

Build when two or more of these are true. Recipes change often enough that document drift is a standing condition rather than an incident. The matrix of jurisdictions and languages has outgrown a folder structure. Your suppliers revise their own sheets and you have no cascade. You cannot produce a list of who holds which version. Or your classification decisions cannot be reproduced from stored inputs when an auditor asks why a product was classified as it was two years ago.

How do hidden costs get into the quote?

Translation is the largest and the most persistently underestimated. The standardised hazard and precautionary statements have official translations, and that part is solved. The rest of a sheet, first aid, firefighting, handling and storage, disposal, is written text that has to be translated accurately and then maintained. That is an ongoing operational cost with a per language, per revision shape, not a one off build line. Ask for it to be modelled at your actual revision rate.

Label printing integration is the second. The label is the artefact a worker actually reads, it is legally sensitive, and it is constrained by physical size, which forces explicit rules about which precautionary statements survive when space runs out. Those rules are a decision your regulatory team must make, and the printing integration itself is fiddly work against specific hardware.

Third is your formulation data audit, discussed above, which is frequently a project before the project. Fourth is licensed content subscription, which is yours to pay whether the software is bought or built and should sit in the total picture rather than appearing later. And fifth is each additional jurisdiction, since each brings its own required content and its own adoption of a particular revision of the globally harmonised system for classification and labelling.

What separates a build that works from one that fails here?

The builds that work treat classification as a derived value that is recomputed whenever any input changes, with the inputs, the rule set version and the result all stored. That gives you two things you cannot otherwise have: a stale list available on demand rather than discovered by a customer, and the ability to answer why a product was classified as it was in March, which requires knowing what was true in March.

They treat the product and its versioned classification as the authoritative object, and every sheet as a rendering produced on demand for a jurisdiction and a language, archived immutably when issued. Manufacturers who instead maintain a folder of authored documents per product get drift multiplied by the number of countries they sell into, which is the failure this whole category exists to prevent.

They govern free text rather than allowing it. Phrases live in a managed library keyed to hazard class and product family, with a translation state, an owner and a review cycle, and authors select rather than type. It is unglamorous and it is what stops the same hazard producing three different first aid paragraphs written by three people in three different years.

And they leave ownership unambiguous. You should own the repository, the infrastructure accounts, your phrase library and the unrestricted right to hire another firm, with licensed regulatory content remaining the licensor's under its own subscription. At Digital Heroes the client owns the code from the first commit. Make sure that distinction is written down, because the phrase library in particular is years of your own regulatory team's work and it should never sit inside somebody else's product.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. The average developer spends more than 17 hours a week dealing with maintenance issues such as debugging and refactoring, and about four of those hours on 'bad code' - waste that equates to nearly $85 billion annually worldwide in opportunity cost. Source: Stripe (2018) →
  2. A 0.1-second improvement in mobile site speed increased retail conversions by 8.4% and average order value by 9.2%; travel conversions rose 10.1%. Source: Deloitte & Google (2020) →
  3. A study (led by Prof. Pak-Lok Poon, published in Frontiers of Computer Science, 2024) reviewing decades of spreadsheet-quality research found that about 94% of spreadsheets used in business decision-making contain errors, illustrating the hidden risk of manual spreadsheet workarounds that custom software is built to replace. Source: Central Queensland University / phys.org (Prof. Pak-Lok Poon et al.) (2024) →
  4. Gartner estimates RPA can eliminate up to 25,000 hours of avoidable rework caused by human errors in the finance function each year, equating to savings of roughly $878,000 for an organization with 40 full-time accounting staff (based on interviews with more than 150 corporate controllers and chief accounting officers). Source: Gartner (2019) →
Prasun Anand · CEO & Founder · New York

Prasun founded Digital Heroes in 2017 and leads it from New York. His work sits where commercial decisions meet delivery: which projects to take on, how teams are shaped across five offices, and where a build is likely to go wrong. Readers get the view from the side that owns the outcome.

View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.

FAQ

Frequently asked questions

Why does a nightly extract of formulations not solve document drift?
Because a product can change and ship on the same day. A batch process picks the change up afterwards, which means the window in which you supply a document describing the previous formulation is exactly the window that causes the problem. The architecture has to be event driven from the system holding the authoritative recipe, so a bill of materials revision immediately marks every affected sheet and label stale and puts it into a work queue somebody owns.
Should we build our own substance and regulatory database?
No. That content represents decades of curation across substance identifiers, harmonised classifications, exposure limits and transport data, plus the ongoing maintenance to keep it current. License it and build the parts that are specific to you: the live recipe connection, your approval workflow, customer specific variants, phrase governance and document generation. Write that boundary into the specification before anyone quotes, including who applies a licensor content update and how.
What happens when a raw material supplier revises their own safety data sheet?
In most manufacturers very little happens until a customer notices. Handled properly, an inbound supplier sheet updates that component's classification, which recomputes the classification of every finished product containing it, which marks the affected sheets and labels stale. Without that cascade your finished product classification silently depends on supplier data that was accurate whenever someone last read it, which for a long tail of materials can be years.
How much of our document archive should we migrate?
Only the issued record. You need the exact document as supplied with its issue date, so a superseded version can be produced if a customer or an auditor questions it, and that is an archive rather than a data set. Do not let anyone parse old documents back into structured classification data, because the result looks authoritative and contains errors you cannot see. Structured data should come from your formulations and licensed content, not from your own prior output.
How do we find out who has an outdated version of a sheet?
You need a distribution register, and if you do not have one the honest answer today is that you cannot. Record every issue of every version to every party, whether that is a customer, a distributor, a portal or one of your own sites, with the date and the channel. Then a revision produces a precise notification list plus evidence that notification happened, which is exactly what a customer audit or a supplier questionnaire asks you to demonstrate.
Our formulations are not all structured in the enterprise system. Is that a blocker?
It is the main schedule risk and it should be audited before the project starts rather than discovered in week six. Count how many finished products have a fully structured composition, how many reference a raw material with no classification on file, and how many are described in prose or reference a supplier blend whose composition you have never held. For each gap decide explicitly whether you request the data, treat the material conservatively, or discontinue the product.
Can the same system produce labels and transport documentation?
It should, because all three derive from the same classification, but they are not the same output. Label content is constrained by physical size, which forces explicit rules about which precautionary statements survive when space runs out, and those rules are a regulatory decision rather than a software one. Transport classification is a separate determination under the applicable road, air and sea rules and can legitimately differ from the supply classification for the same product.
What ongoing costs should we expect after launch?
Translation maintenance is the largest and it scales with your revision rate rather than your product count, because every revised sheet in every language is work. Licensed content subscription continues regardless of whether you buy or build. Each additional jurisdiction adds required content and its own adoption of a particular revision of the globally harmonised system. Budget these as an annual operating line from the start rather than treating the build price as the total cost.
Is it cheaper to customize Salesforce than to build a custom CRM from scratch?
If you use less than a third of what Salesforce does, a custom CRM is often cheaper by year three. Salesforce Enterprise lists at $165 per user per month, so 25 seats cost about $49,500 a year before admin and consultant fees, while a focused custom CRM runs $60,000 to $100,000 once plus 15 to 20% a year in maintenance. If you genuinely need Salesforce's ecosystem, reporting, and app marketplace, customizing it beats rebuilding it; the mistake is paying enterprise prices to use it as a glorified contact list.
How many SaaS seats do we need before building custom becomes cheaper?
The crossover usually shows up between 20 and 50 seats on premium tiers. Salesforce Enterprise lists at $165 per user per month, so 40 users cost about $79,000 a year in subscriptions, which is real money against a custom system you would own outright. Run the comparison over three years: if subscription spend beats the build cost plus 15-20% annual maintenance, custom wins on price before you even count workflow fit.
What is the biggest mistake first-time software buyers make?
Choosing the lowest quote without asking why it is the lowest. A bid 40% under the field usually gets there by skipping tests, documentation, and code review, which are invisible in a demo and brutal to pay for later; every stalled project Digital Heroes has been asked to rescue tells some version of that story. The second mistake is signing without a written scope, which reliably turns the winning cheap quote into 1.5x to 2x the price by launch.
What should I have ready before I contact a development agency?
Three things, none of them technical: a one-page description of the problem in your own words, a list of the tools and spreadsheets the new system must replace or connect to, and a must-have versus nice-to-have split of features. Add a budget range, even a wide one, because it changes the conversation from fantasy to engineering. You do not need a formal specification; producing that is what a discovery phase is for.
What happens to my software if the agency shuts down or we stop working together?
Nothing dramatic, if the engagement was set up correctly: the code sits in your repository, hosting runs on your cloud account, and a handover document explains how to deploy and operate the system. Any competent replacement team can then take over in days rather than months. If the agency controls the repo, the servers, or the domain, fix that now, because renegotiating access during a dispute is the most expensive place to discover the problem.
Does it matter which tech stack the agency wants to use?
Yes, but not in the way most buyers expect: the goal is boring, popular technology such as React, Node.js or Python, and PostgreSQL, because any future team can maintain it and hiring a replacement developer takes days, not months. The red flag is an agency-proprietary framework or an unusual language, which welds you to that one vendor no matter what your contract says about code ownership. A useful test: could you find three freelancers fluent in this stack within a week? If not, push back.
Should I ask for a fixed price or pay the agency hourly?
Fixed price for the first version, hourly or retainer for what comes after launch. A fixed-scope, fixed-price V1 puts the estimation risk on the agency, which is exactly where you want it while trust is unproven; hourly billing on an unscoped greenfield build is a blank check. After launch, flip it, because maintenance and small features arrive unpredictably and fixed-pricing every ticket wastes everyone's time.
How much should a small business budget for its first custom app or website?
For a focused first build, most small businesses land between $8,000 and $60,000: roughly $8,000 to $45,000 for a custom website and $25,000 to $60,000 for an internal tool or simple web app, based on Digital Heroes delivery across 2,000+ projects. Customer-facing products with payments, logins, or a mobile app start around $40,000. Quotes far below these bands usually mean a template with your logo on it, not software shaped around your workflow.
How much should a small business expect to pay for custom software?
Across 2,000+ Digital Heroes projects, a small business system that replaces spreadsheets or one core workflow typically lands between $40,000 and $80,000, with more complex first versions running up to $150,000. The two levers that move the number most are integrations and user roles, not the team's hourly rate. Any quote under $15,000 for a full production system means the vendor has not understood your scope yet.
Should I hire a freelancer or an agency for my software project?
A skilled freelancer is the right call for a single-discipline scope under roughly $15,000, like a website, a plugin, or one integration. Above that, projects need design, backend, testing, and project management at once, and a solo builder becomes the single point of failure: if they get sick or take a bigger client, your project simply stops. Agencies bill 20-40% more per hour but carry continuity, code review, and someone to escalate to, which is what you are actually buying.
How do I vet a software development agency before signing a contract?
Ask to speak with two past clients whose projects resemble yours in size and industry, and ask exactly who will write your code, since some agencies sell senior faces and deliver junior or subcontracted hands. Demand a written specification with acceptance criteria before any fixed price, and check that their portfolio links to products that are actually live. An instant quote given without questions about your workflows is the clearest warning sign there is.
Does the tech stack matter, and which one should I ask for?
It matters less than agencies imply, provided it is boring. A mainstream stack, something like React or Next.js on the front end, Node.js or Python behind it, and PostgreSQL for data, means thousands of developers can maintain your system if you ever change vendors. Apply one test: ask how hard it would be to hire a replacement developer for the proposed stack, and walk away from anything built on an agency's in-house framework.
Who can build a custom software system?

Digital Heroes builds custom software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other software companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading
let's build

Build something worth launching.

A plan, a team, a timeline, within 24 hours. No decks, no discovery calls. Tell us what you're building and we'll come back with a real scope and a real number.

message us directly · we reply within one business day

mission briefing

Monthly dispatch

Playbooks, real build costs, and what we're shipping. One email a month. No fluff.

visit us

New York HQ

1140 Broadway, Suite 704 · New York, NY 10001

Get directions
Online now

Hey there 👋 How can we help you today?