Problems & solutions · Custom Software

Mechanical Integrity Software Problems: The 5 That Cost Real Money, and How to Avoid Them

Mechanical Integrity Inspection Software code editor and API illustration showing common problems and fixes.
The short answer

The most expensive failure is loading legacy thickness readings without first resolving location identity. Do that and the new system produces confident corrosion rates computed from two measurements of two different points, because a marker was painted over during a coating job, a scaffold forced a reading 300 millimetres off, or a spool was replaced and nobody reset the trend. Turnaround scope then gets built on those rates, so you replace piping that did not need replacing and leave a circuit that did, and the second of those two errors is the one that does not stay a budget problem.

Why does loading every unit and every location at once go wrong?

The proposal covers the whole plant: vessels under API 510, piping under API 570, tanks under API 653, every circuit and every condition monitoring location in one migration. It sounds efficient because the data is all sitting in the same spreadsheet estate. What actually happens is that vessels, piping and tanks have genuinely different rules for minimum thickness, interval limits and inspection scope, so the calculation engine has to be right three times before anything is usable once, and the migration work multiplies across three sets of identity problems at the same time.

This is specific to mechanical integrity because the migration is not a load script, it is forensics. Roughly the first third of the effort on a real project is reconciling location identifiers against drawings, resolving duplicates, identifying trends broken by component replacement, and deciding which historical readings are trustworthy enough to keep. Doing that at plant scale before anyone has proven the intake and calculation path means you discover your rules were wrong on ten thousand locations rather than on eight hundred.

The fix is to start with one unit and piping circuits only. Prove the intake pipeline, the validation rules and the rate calculation there, then bring vessels and tanks in behind it with their own rules. In Digital Heroes delivery experience a first release covering the equipment, circuit and monitoring location register, contractor data intake with validation, corrosion rate and remaining life computation with explainable working, and interval scheduling runs $80,000 to $170,000 in 14 to 20 weeks. The full platform adding drawing linkage, risk based inspection support, damage mechanism modelling, repair and temporary repair tracking, turnaround scope generation and mobile field capture runs $220,000 to $500,000 phased over 9 to 18 months.

What goes wrong when spreadsheet thickness history is migrated?

The workbooks come in one per unit, or one per contractor campaign, with location identifiers that only partly match the isometrics. Someone writes a mapping and imports. Three things then go wrong quietly. Readings that were taken at a location before a spool was replaced continue to trend against new metal, producing a negative wall loss or an implausibly low corrosion rate, which is the safest looking wrong answer and therefore the most dangerous. Duplicate identifiers from two campaigns merge into one series that mixes two physical points. And locations that were renumbered when a unit was re rated appear as new series with a single reading each, so no rate can be computed at all and they drop out of the due date report.

What makes this worse than ordinary data migration is that nothing looks broken afterwards. A corrosion rate is a small number, and a wrong small number is indistinguishable from a right one on a dashboard. The error surfaces in a turnaround scoping meeting a year later, as a rate nobody can defend and everyone approves anyway.

The fix is to treat a monitoring location as an object with a history rather than as a key on a reading. It carries its position on the isometric, photographs, access notes, coating and insulation events, and any replacement of the underlying component, and the trend is broken deliberately when metal changes rather than continuing silently. Migrate with a confidence flag on every imported series, review the low confidence ones with an inspector rather than an analyst, and accept that some history is not worth keeping. Budget the forensics explicitly rather than treating it as a load step at the end.

Why do maintenance system and drawing integrations break after launch?

The finding to work order path is demonstrated against SAP Plant Maintenance in a test client and works. In production it starts failing on the cases nobody modelled: a notification type the plant uses for one unit only, functional location codes that were restructured in 2019 with the old codes still present in inspection history, and a work order that gets rejected because a required field is populated differently by the reliability group than by inspection. Drawing linkage degrades a different way. It works for the intelligent isometrics and quietly does nothing for the scanned paper drawings, which is most of the older units.

The difficulty is that your equipment register in SAP PM or Maximo was built for maintenance while your circuit and location structure was built for integrity, and the two naming conventions were never reconciled. Every mismatch becomes either a manual step or a wrong link.

The fix is to make the mapping between the maintenance hierarchy and the integrity hierarchy an explicit, maintained object with an owner, not an assumption buried in an interface. Fail loudly: a finding that cannot create a notification should sit in a queue with the reason, not vanish. For drawings, separate the two cases in the plan and price them separately, because placing locations on scanned paper isometrics is slow manual work and pretending otherwise is how the schedule slips. Ask a developer what they have integrated by name, at which plant and which version, because SAP PM notifications, Maximo work orders and an intelligent isometric system are four separate problems.

What happens when API interval rules are not covered properly?

The system computes a corrosion rate, divides remaining wall by it, and produces a next inspection date. That is arithmetic, not compliance. API 510, 570 and 653 each relate inspection intervals to remaining life and corrosion rate in their own way, they distinguish short term from long term rates for good reasons, and they cap intervals regardless of what the calculation returns. A build that implements the division and not the caps will hand an inspector a date that is longer than the standard allows, and the inspector will not notice on every circuit.

Then there are the awkward cases, which are most of them at scale. A location with only one reading. A location where the last two readings disagree with the ten year trend. A component replaced mid history. A grid where the governing reading moves from one point to another between campaigns. Wall loss over an interval that sits close to the accuracy of the instrument, so the computed rate is mostly measurement uncertainty.

The fix is to compute both short term and long term rates, apply your own documented governing rules, calculate remaining life against the correct minimum thickness for that component, and cap the resulting date at the maximum interval in the applicable standard. Then make every number explainable: click the rate and see the readings behind it, the rule applied and anything excluded. An integrity system that cannot show its working is one an auditor will not accept and an engineer should not sign, and no amount of interface design compensates for that.

Should you build custom or configure what you already own?

Buy Metegrity Visions, Cenosco IMS or Antea if you are prepared to adapt your practice to theirs. These are mature products with sound data models, used at serious plants, and there is no shame in choosing one, particularly where a corporate group has already standardised. Buy GE Vernova APM if you are already inside that ecosystem for reliability and want integrity in the same place. If you run a small terminal with a few hundred monitoring locations and one contractor, a disciplined spreadsheet with a competent inspector is proportionate and much cheaper than either.

Build when two or more of these are true. Your inspection data is spread across spreadsheets with location identifiers that do not reliably match your drawings. Your contractors return data in formats a package cannot ingest without manual rework each campaign. You have implemented a package and are running a parallel spreadsheet for the calculations your engineers actually trust. Your risk based inspection study is a static report while process conditions have moved. Your turnaround scope cannot be traced line by line back to the data that justified it.

The threshold is scale multiplied by data quality. Above roughly a few thousand locations with several contractors and a drawing estate accumulated over decades, the coordination problem has become a safety problem, and that belongs in a system you own and can correct the week you find a fault in it. If you are already paying for a package and still keeping the spreadsheet, that is the signal.

How do hidden costs get into the quote?

Five places, and four of them are data rather than code.

  • Existing data condition. The dominant driver. A quote that treats migration as an import task has not looked at your workbooks. Ask for it to be priced as a reviewed forensic exercise with a stated sample already inspected.
  • Drawing linkage. Intelligent isometrics and scanned paper are different jobs. Price them separately with a count of each, or the paper ones will consume the contingency.
  • Contractor formats. Each additional contractor is a mapping, cheap once the intake pipeline exists and not free before it. List them by name in the scope.
  • Risk based inspection depth. A qualitative model tied to inspection history costs a fraction of a full quantitative implementation. A quote that says risk based inspection without saying which one is hiding a large number.
  • Maintenance system integration. Notifications and work orders flowing without double entry is real work, and it depends on a hierarchy mapping somebody has to own after go live.

What separates an integrity build that works from one that fails?

Four things, all testable in a first conversation.

Ask what happens when a spool is replaced. If readings simply continue to trend against the new metal, the system will generate corrosion rates that are nonsense in the safest looking direction, and that is the failure mode that matters most. The correct answer involves breaking the trend deliberately and recording the component change as an event on the location.

Ask how a corrosion rate will be explained. Every number must open into the readings behind it, the rule applied and anything excluded. Engineers will not stake a remaining life assessment on a figure they cannot audit, and if they cannot audit it they will keep the parallel spreadsheet, which means you have bought a second system rather than replaced the first.

Ask how contractor data is validated on arrival. If the answer is import and review, the same transcription and identity errors you have today will simply arrive faster. Quarantine rules are the minimum: a reading above the previous one beyond measurement uncertainty, a physically implausible implied rate, an unknown location, a gap where a grid point was read last time, each held until accepted with a recorded reason.

Settle ownership in writing before kickoff: the repository, the infrastructure accounts and the right to hire another firm. At Digital Heroes the code is yours from the first commit. Integrity data is the evidence that your pressure envelope is fit for service, and it should never live somewhere you cannot retrieve it in full.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. Across 1,471 IT projects the average cost overrun was 27%, but one in six projects was a 'black swan' with an average cost overrun of 200% and a schedule overrun of nearly 70%. Source: Harvard Business Review (Bent Flyvbjerg & Alexander Budzier, University of Oxford) (2011) →
  2. A 100-millisecond delay in website load time can cut conversion rates by 7%; a two-second delay increases bounce rates by 103%; and 53% of mobile visitors leave a page that takes longer than three seconds to load. Source: Akamai Technologies (2017) →
  3. 88% of customers say good customer service makes them more likely to purchase from a brand again in the future, quantifying the direct revenue link between support quality and retention. Source: HubSpot (2024) →
  4. An analysis of enrollment and completion data for 221 MOOCs (Katy Jordan, published in the International Review of Research in Open and Distributed Learning, IRRODL, 16(3), 2015 - not the Journal of Distance Education) found completion rates ranging from 0.7% to 52.1%, with a median completion rate of 12.6%, and completion negatively correlated with course length (longer courses had lower completion rates) - underscoring how unsupported self-paced online courses struggle to finish learners. Source: Journal of Distance Education (via ERIC / Katharina Jordan) (2015) →
Aanya B. · Senior Frontend Engineer · Next.js · Delhi

Aanya builds frontends in Next.js at Digital Heroes, covering rendering strategy, component structure, accessibility and the performance work that decides how a site feels on a mid range phone. Her writing translates frontend decisions into the outcomes non technical stakeholders actually care about.

View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.

FAQ

Frequently asked questions

Our engineers still keep a parallel spreadsheet after implementing a package. Is that a data problem or a software problem?
It is usually a trust problem with a specific cause: the package computes a rate they cannot open and audit, or it applies a governing rule that differs from the one their procedure specifies. Find the exact circuit where the two disagree and work backwards, because that single case will tell you whether the fix is configuration, a mapping error in migration, or a genuine gap. A parallel spreadsheet is a symptom worth diagnosing rather than a habit to discourage.
How do we deal with a reading that came back higher than the last one?
Treat it as a quarantine event, not a data point. A wall thickness that increased is either a measurement taken at a different point, an instrument or calibration issue, a transcription error, or a component that was replaced without the trend being reset. All four need a human decision, and all four are cheap to resolve at intake and expensive to resolve a year later when the rate has already fed a turnaround scope.
Can we keep our existing risk based inspection study?
Yes, and you should, but hold it as live data rather than as a document. Damage mechanisms per circuit, consequence category, inspection effectiveness and the resulting interval belong in the system so that when a feedstock or process condition changes, the affected circuits are identifiable immediately rather than at the next study. That does not replace the engineering judgement in a formal API 580 or 581 assessment. It stops that judgement decaying into a PDF nobody revises.
How should temporary repairs such as clamps be tracked?
As first class objects with an expiry date, an owner and a mandatory review, not as a note in an inspection report. The failure mode is well known: a clamp installed under an approved temporary repair procedure has a defined life and a plan to replace it, and the plan surfaces only when somebody remembers. Making the expiry a tracked item with escalation is one of the cheapest safety improvements in this category.
What do we do about monitoring locations that only have one reading?
Report them as a distinct category rather than letting them fall out of the due date list, because a location with no computable rate is not a location with no risk. Options are to schedule a reading to establish a rate, to apply a conservative rate from a comparable circuit with the assumption recorded, or to flag the location for review. What matters is that the choice is explicit and visible, since silently excluded locations are how gaps persist for years.
How long does migration from spreadsheets really take?
Expect roughly the first third of the project effort, and expect it to be the honest reason schedules slip. The work is reconciling identifiers against drawings, resolving duplicates, finding trends broken by component replacement and deciding which history is trustworthy. Budget it as its own workstream with an inspector involved rather than an analyst alone, because most of the decisions require someone who knows the plant.
Can inspection findings generate turnaround scope automatically?
They can, and it changes scoping meetings from opinion to evidence. Findings become tracked items with owners and required action dates, and scope is generated from open items with each line traceable to the thickness trend or inspection report that justified it. The uncomfortable and useful consequence is that scope which cannot be justified from data becomes visible for what it is, which is a conversation worth having before the shutdown rather than during it.
Should the field capture app work offline?
Yes, because scaffolded areas, tank interiors and older units frequently have no usable signal, and an inspector who cannot record at the point of measurement will record later from a notebook. That reintroduces exactly the transcription errors the validation pipeline exists to catch. Local capture with validation applied on the device and queued sync when signal returns is the pattern that holds up in a plant environment.
How much should a small business budget for its first custom app or website?
For a focused first build, most small businesses land between $8,000 and $60,000: roughly $8,000 to $45,000 for a custom website and $25,000 to $60,000 for an internal tool or simple web app, based on Digital Heroes delivery across 2,000+ projects. Customer-facing products with payments, logins, or a mobile app start around $40,000. Quotes far below these bands usually mean a template with your logo on it, not software shaped around your workflow.
Couldn't I just build my app in Bubble or another no-code tool instead of hiring an agency?
For validating an idea with real users, yes, and we tell clients that honestly. The walls come later: Bubble apps cannot be exported as code to run anywhere else, performance drops on complex data operations, and usage-based pricing climbs as you grow. A meaningful share of Digital Heroes custom builds are rebuilds of no-code MVPs that proved the business worked, which is the system operating as intended: validate cheap, then build the version that scales.
We run everything on Airtable and spreadsheets. When is it time to go custom?
The switch usually makes sense when you hit one of two walls: Airtable's record caps (125,000 records per base on the Business plan) or logic the tool cannot express, like multi-step approvals with conditional pricing. There is also a simple cost signal: 25 people on Business at roughly $45 per seat per month is about $13,500 a year, forever, for a tool you are already fighting. Custom is worth it when the workflow is core to how you make money; for peripheral processes, staying on Airtable is the right call.
What does it cost to keep custom software running after launch?
Budget 15-20% of the original build cost per year, which on a $100,000 system means $15,000 to $20,000 for security patches, dependency updates, bug fixes, and small improvements as real usage reveals what the spec missed. Cloud hosting for a typical business application adds $50 to $300 a month on top. Skipping maintenance does not save the money; in Digital Heroes rescue work, unmaintained systems typically need a far more expensive rebuild within about three years.
How do I make sure custom software is secure and compliant with rules like HIPAA?
Start with the baseline every business system should have: encryption in transit and at rest, role-based access control, and audit logs. If HIPAA applies, the hosting provider must sign a Business Associate Agreement, which AWS, Azure, and Google Cloud all offer, and access controls have to be designed in from day one, not bolted on. SOC 2 certifies a company's operating practices, not a codebase, so ask vendors what they have shipped in your regulated domain rather than which logos are on their website.
Does it matter which tech stack the agency wants to use?
Yes, but not in the way most buyers expect: the goal is boring, popular technology such as React, Node.js or Python, and PostgreSQL, because any future team can maintain it and hiring a replacement developer takes days, not months. The red flag is an agency-proprietary framework or an unusual language, which welds you to that one vendor no matter what your contract says about code ownership. A useful test: could you find three freelancers fluent in this stack within a week? If not, push back.
What should I have ready before I contact a development agency?
Three things, none of them technical: a one-page description of the problem in your own words, a list of the tools and spreadsheets the new system must replace or connect to, and a must-have versus nice-to-have split of features. Add a budget range, even a wide one, because it changes the conversation from fantasy to engineering. You do not need a formal specification; producing that is what a discovery phase is for.
How do I vet a software development agency before signing a contract?
Ask to speak with two past clients whose projects resemble yours in size and industry, and ask exactly who will write your code, since some agencies sell senior faces and deliver junior or subcontracted hands. Demand a written specification with acceptance criteria before any fixed price, and check that their portfolio links to products that are actually live. An instant quote given without questions about your workflows is the clearest warning sign there is.
If an agency builds my software, who actually owns the code?
You should own everything, assigned in writing: the contract transfers full IP to you on final payment, the code lives in your GitHub organization, and hosting runs in cloud accounts you control. The red flag is a proposal that mentions the agency's proprietary platform or framework, which usually means you are renting, not buying. Digital Heroes structures every build this way precisely so a client can fire us and lose nothing but the relationship.
What is a discovery phase, and is it worth paying for separately?
Pay for it, and treat the output as yours. A discovery phase runs two to three weeks, typically 5 to 10% of the eventual build budget, and produces a written scope, wireframes, and a fixed quote you can take to any vendor, including a competitor of the agency that wrote it. Skipping it is how projects end up quoted from a two-paragraph email and delivered at twice the price.
Should I ask for a fixed price or pay the agency hourly?
Fixed price for the first version, hourly or retainer for what comes after launch. A fixed-scope, fixed-price V1 puts the estimation risk on the agency, which is exactly where you want it while trust is unproven; hourly billing on an unscoped greenfield build is a blank check. After launch, flip it, because maintenance and small features arrive unpredictably and fixed-pricing every ticket wastes everyone's time.
Who can build a custom software system?

Digital Heroes builds custom software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other software companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading
let's build

Build something worth launching.

A plan, a team, a timeline, within 24 hours. No decks, no discovery calls. Tell us what you're building and we'll come back with a real scope and a real number.

message us directly · we reply within one business day

mission briefing

Monthly dispatch

Playbooks, real build costs, and what we're shipping. One email a month. No fluff.

visit us

New York HQ

1140 Broadway, Suite 704 · New York, NY 10001

Get directions
Online now

Hey there 👋 How can we help you today?