Problems & solutions · Custom Software

Process Safety Management Software Problems: The 7 That Cost Real Money, and How to Avoid Them

Process Safety Management Software workflow illustration showing common problems and fixes.
The short answer

The most expensive failure in this category is importing hazard studies as attachments. It is quick, it looks like progress, every study is now searchable in one place, and it changes nothing about the question that actually matters. When an auditor asks whether the high high level trip credited at node 12 is still installed, still set at the value in the study and still proof tested at its required interval, the answer is still forty minutes of screen sharing across three systems ending in an uncertain answer. You have bought a document library with a login, at the price of a safety system, and the underlying exposure, which is that credited protection may not exist, is exactly where it was before the project started.

Why does the plan to replace the facilitation tool keep going wrong?

The first workshop tends to conclude that the study tool is the problem, because it is the tool everyone touches. Facilitators complain about it, the reports are ugly, the export is awkward, and a replacement feels like the obvious scope.

It is the wrong target. Sphera PHA-Pro is a facilitation environment and it does that job well: it captures a hazard and operability study or a layer of protection analysis efficiently in a room with a team and a facilitator, and it produces the report. Rebuilding that means competing with a mature product on its strongest ground, using budget that was raised to solve a different problem.

The real failure sits after the study finishes. The output leaves the tool and becomes a document, and the safeguards it credits stop being tracked objects the moment the report is issued. Nothing connects them to the equipment they describe, the maintenance regime that tests them, the changes that alter them or the bypasses that suspend them. That is where the trail goes cold, and no facilitation tool was ever supposed to fix it.

There is also a practical cost to replacing facilitation. Facilitators are experienced people with strong preferences, and forcing a new study tool on them mid-programme slows studies down at exactly the moment you want their goodwill for the register work.

The fix is to keep the facilitation tool if your teams like it, and build the register and the links around it. Import studies as structured objects, meaning nodes, deviations, causes, consequences, safeguards, risk rankings and recommendations with their relationships intact, and let the room keep working the way it works. Replacing a facilitation tool is rarely the win. Owning a live safeguard register always is.

What goes wrong when you extract safeguards from legacy studies?

This is the hardest part of the project and it is routinely priced as a data load.

The problem is that a credited safeguard in a 2004 study is a sentence written by a facilitator under time pressure. It reads as a high high level trip on the flash drum, or a relief valve on the vessel, or operator response to alarm. None of those are identifiers. Turning them into references to a real instrument tag, relief device, procedure revision or interlock requires someone who knows the plant, and the plant has changed since.

What emerges from that exercise is uncomfortable in a useful way. Safeguards credited against equipment removed in a revamp. Two studies crediting the same device for different scenarios with different assumptions about its availability. Procedures cited by a title that no longer exists. Sites consistently find things they did not expect, which is precisely why the exercise never gets started voluntarily.

The failure mode in the project is either scope or false confidence. Scope failure is trying to resolve every safeguard in every study before anything goes live, which stalls the build behind a process safety engineer who has a day job. False confidence is worse: accepting machine proposed tag references without adjudication, producing a register that looks authoritative and is partly fiction. An unadjudicated register is worse than no register, because it invites reliance.

The fix is a queue rather than a project. Machine assistance reads the legacy worksheets and proposes the tag each phrase implies, and a process safety engineer confirms or corrects every single one. Anything unresolvable goes on an exception list that is itself a deliverable, because that list is where the genuine surprises live. Sequence by covered process, highest risk first, and let the register go live partially resolved with resolution status visible on every entry.

Why do the maintenance and control system integrations break after launch?

Two integrations carry the whole value of this system, and both fail in ways that are silent.

The maintenance integration breaks on numbering. Your studies reference equipment tags from the instrument index and process drawings. Your computerised maintenance system frequently carries a different identifier, invented when the system was implemented, sometimes with a prefix or a location code baked in. The mapping between them is often a spreadsheet held by one reliability engineer. When it is wrong or incomplete, safeguard test status simply does not populate for a subset of items, and a blank looks like a system still loading rather than a gap.

The second maintenance failure is job scope. A preventive maintenance job against a tag is not proof of a proof test. The job may cover a calibration rather than a functional test of the trip, and a system that treats any completed job on that tag as evidence of testing will report green on a safeguard nobody has functionally tested in three years.

The control system integration breaks on the network boundary. Reading bypass status directly is the right answer, and it crosses from the control network into the business network, which means the controls engineer, a security review and an architecture that reads without ever writing. Projects that leave this to the end discover the security review has a lead time measured in months.

The fixes are unglamorous. Build and own the tag crosswalk as data with a completeness measure reported on a dashboard, so unmapped safeguards are visible rather than blank. Match on job type, not just on tag, so only the job that constitutes a proof test counts as one. And start the control network conversation in week one, with a read-only pattern proposed before anyone asks.

What happens when bypass control and revalidation are not covered?

Two obligations get deferred to phase two and both undermine the register they were meant to support.

Bypasses are the sharper one. A safety instrumented function credited in a layer of protection analysis is only protective when it is in service, and the integrity level assumed in the calculation assumes availability. Every hour the function spends bypassed is an hour the risk assessment does not describe. A register that shows a safeguard as present, tested and healthy while it is physically bypassed is not neutral, it is actively misleading, and it is worse than the spreadsheet it replaced because people trust it.

The specific failure is that bypasses are recorded by shift teams in an operations log the safety system never reads. So the register and the plant disagree, and the disagreement is invisible until an incident investigation finds it.

Revalidation is the slower failure. The regulation requires hazard analyses to be revalidated at least every five years, and if the previous study exists only as a report, the team starts by retyping node structures and re-debating scenarios that were settled a decade ago. Facilitated time is expensive and it gets spent on transcription rather than judgement.

The fix is to treat a bypass on a credited safeguard as an event that changes the site's risk position, with duration limits related to integrity level, compensating measures recorded against a named person, automatic escalation as the limit approaches, and a shift handover review where the incoming crew accepts a specific list. And to hold studies as structured data from the start, so revalidation begins from the previous study with modifications, incidents, closed and deferred recommendations and poor test histories already highlighted per node.

Should you build custom or configure what you already own?

Some sites should not build, and it is worth saying which.

If you operate one covered process with one current study and a recommendation list you can read in a single sitting, a well governed spreadsheet plus a facilitation tool is proportionate. The value of this category comes from scale and from linkage, and at that size you have neither problem.

If your corporation has already standardised on Enablon, Intelex or VelocityEHS for incidents and audits, your sites genuinely share a risk matrix, and your ambition is action tracking rather than live safeguard status, use the module you already own. Action tracking in those platforms is real and it is more than a spreadsheet.

Before commissioning anything, find out what your existing platform already does and is not configured for. Recommendation workflows with risk based due dates and evidence based closure are frequently available and switched off, usually because the implementation was done years ago by a corporate team with different priorities.

The build case starts when two or more of these are true. You cannot currently produce a list of every safeguard your site credits. Your studies reference equipment that has been modified since and nobody has traced the impact. You have safety instrumented functions whose proof test intervals need to be visibly linked to what they were credited for. Your bypass register and the actual plant have disagreed at least once. Or you run multiple sites with different node structures and risk matrices, where a single corporate template would degrade all of them to make one report tidy.

The honest test is whether you are managing documents or managing risk. Demonstrating that studies were completed is document management and a platform will do. Answering at any moment whether the protection you claim is in place and tested is a live register, and a live register has to know your tags.

How do hidden costs get into the quote?

The application build is usually estimated fairly. These items arrive afterwards.

  • Legacy study condition. A recent study exported cleanly is a different problem from a 2004 worksheet in a scanned document. Count your studies by vintage and format before anyone estimates.
  • Process safety engineer time. Adjudicating proposed safeguard-to-tag resolutions is the largest human cost in the project, it cannot be delegated outside the site, and it competes with a full workload.
  • The tag crosswalk. If your maintenance system uses different identifiers from your instrument index, building and validating that mapping is a workstream owned by reliability engineering.
  • Control network access. Reading bypass status crosses a security boundary. Budget the review lead time, not just the engineering.
  • Multiple sites. Different risk matrices and node conventions mean you are supporting several models rather than configuring one, and forcing convergence is a change management project in its own right.
  • Safety instrumented function reliability data. If in scope, it brings its own calculation, its own audit expectations and its own specialist review.

What separates a build that works from one that fails here?

Four things, and the first is a question you should ask in the first meeting.

Ask a prospective developer to explain how a safeguard becomes a tag. Someone who has done this work will describe extraction, adjudication by a named engineer, an exception list for safeguards that cannot be resolved, and a plan for the ones that turn out to reference removed equipment. Someone who proposes importing studies as attachments has understood the filing problem and missed the safety problem entirely.

The second is that a change record finds the studies it affects mechanically. When a change touches equipment tags, the system should list every hazard study node and every credited safeguard referencing those tags, so the engineer applies judgement to a complete list rather than to a search. If the answer is that an engineer selects the affected studies from a dropdown, nothing that mattered has been automated. This link is the most consequential integration in the category and it is almost never in place.

The third is that closure requires evidence of the right kind. A marked up drawing, a revised procedure with a revision number, a completed work order, a training record. Recommendations closed with a comment saying done are not closed, and a system that accepts that is a filing cabinet with a login. Deferring a high consequence recommendation should require a named senior signature and a documented interim measure, and that interim measure becomes a safeguard in the register with its own expiry.

The fourth is ownership. Own the repository, the infrastructure accounts and the right to hire anyone else, in writing before kickoff. At Digital Heroes the code is yours from the first commit. A safeguard register is evidence in a regulatory inspection and after an incident, and evidence should never depend on somebody else's licence terms remaining current.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. Only 22% of firms are 'future ready' having significantly transformed digitally; these companies show average revenue growth 17.3 percentage points and net margins 14.0 percentage points above their industry average. Source: MIT Center for Information Systems Research (MIT Sloan) (2022) →
  2. McKinsey found that tech debt can amount to 20-40% of the value of a company's entire technology estate before depreciation, and CIOs report that 10-20% of the budget for new products is diverted to resolving tech-debt issues. Source: McKinsey & Company (2020) →
  3. Grand View Research valued the global field service management market at USD 4.43 billion in 2022 and projects it to reach USD 11.78 billion by 2030, a 13.3% CAGR, driven by growing field operations in telecom, utilities, construction and energy. Source: Grand View Research (2023) →
  4. The average number of formal learning hours used per employee fell to 13.7 in 2024, down from 17.4 in 2023, a decline the report attributes partly to a shift toward informal and on-the-job learning not captured in the formal-hours metric. Source: Association for Talent Development (ATD) (2025) →
Layla S. · Senior Account Manager · Wellness · Sydney

Layla looks after wellness sector accounts, running projects that touch bookings, memberships, subscriptions and the customer data that sits behind them. She translates between clinical or operational language and what a development team needs written down. Useful reading if your business runs on recurring relationships rather than one off sales.

View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.

FAQ

Frequently asked questions

Should we replace PHA-Pro as part of this?

Usually not. It is a good facilitation environment for running the study in a room with a team, and replacing it means competing with a mature product on its strongest ground while the actual problem goes untouched. The gap is what happens after the study finishes, when the output becomes a document and the safeguards it credits stop being tracked objects. Keep the facilitation tool, import studies as structured data, and build the register and the links around it.

How long does resolving safeguards to equipment tags take?

Longer than the software, because the constraint is a process safety engineer with a day job adjudicating every proposal. Machine assistance reads legacy worksheets and proposes which tag a phrase implies, but a register nobody has adjudicated is worse than no register since it invites reliance. Sequence by covered process with the highest risk first, let the register go live partially resolved with resolution status visible, and treat the unresolvable list as a deliverable in its own right.

Why does safeguard test status not populate for some equipment?

Almost always a tag mismatch. Your studies reference identifiers from the instrument index and drawings while your maintenance system carries its own numbering invented at implementation, and the mapping between them is often a spreadsheet held by one reliability engineer. Own the crosswalk as data with a completeness measure on the dashboard, so unmapped safeguards appear as a known gap rather than as a blank that looks like a screen still loading.

Is a completed maintenance job proof that a safeguard was tested?

No, and treating it that way is a quiet and serious failure. A preventive job against a tag may cover a calibration rather than a functional test of the trip, so a system that counts any completed job as evidence will report a healthy safeguard that has not been functionally tested in years. Match on job type as well as tag, so only the job that constitutes a proof test counts as one.

What is the risk of leaving bypass tracking to phase two?

The register shows a safeguard as present, tested and healthy while it is physically bypassed, which is worse than the spreadsheet it replaced because people trust it. Bypasses are typically recorded by shift teams in an operations log the safety system never reads, so the register and the plant disagree invisibly. Treat a bypass on a credited safeguard as a change to the site's risk position with duration limits, verified compensating measures and shift handover acceptance.

What has to happen before we can read bypass status from the control system?

A conversation with the controls engineer and a security review, because the read crosses from the control network into the business network. The right architecture reads and never writes, and it should be proposed before anyone asks for it. Start this in week one rather than at the end, since the review lead time is frequently measured in months and it is the item most likely to delay a phase that is otherwise complete.

How should recommendation closure be controlled?

By requiring evidence of the appropriate kind: a marked up drawing, a revised procedure with a revision number, a completed work order or a training record. Closure by comment is not closure. Recommendations should inherit the risk of the scenario they address, which sets the due date and the approval level required to defer them, and deferring a high consequence item should need a named senior signature plus a documented interim measure that enters the register with its own expiry.

Which costs are usually missing from the estimate?

Process safety engineer time for adjudication, which is the largest human cost and cannot be sourced externally. The condition of legacy studies, since a scanned 2004 worksheet is a different problem from a recent clean export. Building and validating the tag crosswalk. Security review lead time for control network access. Supporting multiple sites with different risk matrices and node conventions. And safety instrumented function reliability data if it is in scope.

Who owns the code when an agency builds my software?
You should, completely, through a written intellectual property assignment that transfers everything on final payment; without that clause, copyright stays with whoever wrote the code by default. Insist that the repository lives in your own GitHub organization from day one and that hosting, domains, and third-party accounts are registered to you. Also check for licenses to the agency's proprietary frameworks buried in the contract, because those can make switching vendors practically impossible even when you own your own code.
Should I hire a freelancer or an agency for my software project?
A skilled freelancer is the right call for a single-discipline scope under roughly $15,000, like a website, a plugin, or one integration. Above that, projects need design, backend, testing, and project management at once, and a solo builder becomes the single point of failure: if they get sick or take a bigger client, your project simply stops. Agencies bill 20-40% more per hour but carry continuity, code review, and someone to escalate to, which is what you are actually buying.
If an agency builds my software, who actually owns the code?
You should own everything, assigned in writing: the contract transfers full IP to you on final payment, the code lives in your GitHub organization, and hosting runs in cloud accounts you control. The red flag is a proposal that mentions the agency's proprietary platform or framework, which usually means you are renting, not buying. Digital Heroes structures every build this way precisely so a client can fire us and lose nothing but the relationship.
What should I have ready before I contact a development agency?
Three things, none of them technical: a one-page description of the problem in your own words, a list of the tools and spreadsheets the new system must replace or connect to, and a must-have versus nice-to-have split of features. Add a budget range, even a wide one, because it changes the conversation from fantasy to engineering. You do not need a formal specification; producing that is what a discovery phase is for.
What should I prepare before contacting a software development agency?
A one-page brief beats a 40-page requirements document: the business problem in plain words, who will use the system, the 5 to 10 workflows it must handle, the tools it must connect to, and your budget range and deadline driver. You do not need wireframes, a specification, or technical vocabulary; producing those is the agency's job during discovery. Stating a budget range up front is the single best move, because it gets you honest scoping instead of a quote engineered to win the meeting.
How long does it take from first call to software my team can actually use?
Plan for four to six months: two to three weeks of discovery, two to four weeks of design, then a 10 to 16 week build with testing. In Digital Heroes delivery experience the schedule killer is not engineering speed but decision lag; a client who takes two weeks to approve wireframes adds two weeks to launch. Book a weekly 30-minute decision slot before kickoff and most of that risk disappears.
Couldn't I just build my app in Bubble or another no-code tool instead of hiring an agency?
For validating an idea with real users, yes, and we tell clients that honestly. The walls come later: Bubble apps cannot be exported as code to run anywhere else, performance drops on complex data operations, and usage-based pricing climbs as you grow. A meaningful share of Digital Heroes custom builds are rebuilds of no-code MVPs that proved the business worked, which is the system operating as intended: validate cheap, then build the version that scales.
Will custom software work with the tools we already use, like QuickBooks and Stripe?
Yes, and this is one of custom software's genuine advantages: QuickBooks, Stripe, Shopify, and most mainstream business tools publish documented APIs built for exactly this. Expect each standard integration to add one to two weeks of build time, and be suspicious of any quote that lists five integrations without asking what data flows in which direction. The hard cases are legacy systems with no API, which is a question to raise in discovery, not in week nine.
If we build for 20 users now, will the software cope with 500 later?
It should, without a rewrite, if it was built on a standard cloud stack; going from 20 to 500 users is mostly a hosting configuration change costing hundreds a month, not a second project. What actually breaks under growth is sloppier work: database queries never indexed for volume and features designed assuming one office's worth of data. Before signing, ask the vendor what happens to the system at ten times today's data, and listen for a specific answer.
Who can build a custom software system?

Digital Heroes builds custom software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other software companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading
let's build

Build something worth launching.

A plan, a team, a timeline, within 24 hours. No decks, no discovery calls. Tell us what you're building and we'll come back with a real scope and a real number.

message us directly · we reply within one business day

mission briefing

Monthly dispatch

Playbooks, real build costs, and what we're shipping. One email a month. No fluff.

visit us

New York HQ

1140 Broadway, Suite 704 · New York, NY 10001

Get directions
Online now

Hey there 👋 How can we help you today?