Dam and Levee Safety Monitoring Software Problems: The 7 That Delay an Alert, and How to Avoid Them
The most expensive failure is the interval between a reading being taken and an engineer seeing it. A technician reads a standpipe piezometer on a Tuesday walk, writes it on a clipboard, photographs the sheet that evening and emails it Thursday, and the value only becomes visible when somebody opens the spreadsheet. If that reading crossed a threshold, the structure has been in an unreviewed condition for days, and whether anyone acted depended on one engineer's file habits rather than on any control. Trace that interval for your last exceedance and you have measured both the risk and the business case.
Why does trying to bring the whole portfolio in at once fail so often?
The biggest scope failure in dam and levee monitoring builds is starting with every structure and every instrument. An owner with eleven structures assembles a specification covering two concrete gravity dams on dataloggers, six earth embankments read by hand, two levee reaches on survey monuments and a tailings facility under a different regulatory regime, and expects one release to cover all of it.
The instrument diversity alone defeats that. A portfolio assembled over four decades carries several manufacturers, several datalogger generations, and some instruments whose documentation no longer exists. Each brings its own calibration handling and its own quirks, and each is a discovery exercise with the engineer who knows the structure. Attempt them in parallel and the project spends months in data archaeology with nothing running.
Start with the structures that carry the highest hazard classification and the population at risk, get their readings, conversions, thresholds and alerting working, and extend afterwards. A first release covering unified capture of automated and manual readings, versioned instrument metadata with correct conversions, governed thresholds and acknowledged alerting runs $60,000 to $120,000 over 10 to 14 weeks in our delivery experience. That release changes daily practice at the structures where it matters most, and it teaches the team what the metadata model has to survive before the awkward instruments arrive.
What goes wrong with historical readings, calibration and datums?
The forty year piezometric record is the asset, and importing it is where these projects most often do quiet damage.
The first problem is conversion history. Vibrating wire piezometers need their calibration polynomial and temperature correction applied, barometric compensation matters for some installations, and everything has to resolve to a piezometric elevation against a surveyed reference. Those constants change when an instrument is recalibrated and the reference changes when a monument is re-surveyed. If the import applies today's constants to the whole series, every historical reading is silently restated and the trend the engineer has relied on for a decade shifts. Nobody notices, because the new numbers look plausible.
The second is reference datums. Spreadsheets maintained since the 2000s frequently mix elevations from different survey epochs, sometimes within the same file, and a portfolio comparison across structures on different datums is not comparison at all.
The third is data quality that was handled by human memory. The engineer knows which instrument has been unreliable since a lightning strike and which one runs high after rain, and the spreadsheet contains those readings with no marker. Imported wholesale, they become evidence.
Version instrument metadata with effective dates so past readings keep the conversion that applied when they were taken. Resolve datums explicitly and record which epoch each elevation came from. And spend interview time with the engineer flagging known suspect periods as declared exclusions before the import rather than after. Many owners migrate the last ten years first, prove the model, then backfill older history once the metadata handling has met real edge cases.
Why do the datalogger and survey integrations break after launch?
Field integrations in this domain break for physical reasons more than software ones, which means they break at the worst times, typically during the wet season when the data matters most.
Telemetry links drop. A remote site loses power, a modem fails, a datalogger program is edited during a site visit and the export format changes. The failure mode that matters is not the outage, it is the silence: nothing alerts because the last successful load looked normal, and three weeks later somebody notices a gap. Every instrument needs an expected reading frequency and an alert when readings stop arriving, which is a different and more important alert than a threshold crossing.
Manual capture breaks differently. A field application that needs connectivity to record a reading will be abandoned in week two, because remote embankments and levee reaches have no signal and a technician is not going to stand in a field waiting. Offline capture with local storage, the previous value displayed at the point of entry so an implausible number is caught on site, photograph attachment and reliable sync afterwards is a release one requirement rather than an enhancement.
Survey data arrives as consultant reports on their own schedule and in their own format, sometimes as documents rather than data. Treat that ingestion as a defined interface with the consultant, agreed in the appointment, or it becomes a manual retyping task that nobody sustains. The same applies to any third party who supplies readings on your behalf.
What happens when threshold governance and recommendation closure are not covered?
Two gaps get deferred routinely, and both show up at inspection rather than in the data.
Threshold governance is the first. In most portfolios, thresholds live inside datalogger programs, set at commissioning by whoever configured the station, and nobody can produce the rationale for their current values. That is a finding waiting to happen, because a threshold on a safety instrument is an engineering judgement and should carry the name of the engineer of record who set it, the date, the reasoning and a revision history. Held as data with controlled change, it also becomes reviewable: an engineer can look at every threshold in the portfolio in one place, which is not possible when they are scattered across station programs.
Prior recommendation closure is the second and it is the one that most often goes wrong. Whether your reviewer is a federal energy regulator requiring independent consultant inspections on a cycle, a state dam safety office with its own format, or a mining authority under a tailings standard, the next inspection will ask about every recommendation raised in the last one. Owners track them in a spreadsheet or not at all, and an owner who cannot show closure evidence for a recommendation issued five years ago is having a much longer conversation than the underlying data warranted.
Make recommendations tracked objects with owners, due dates, closure evidence and a link to the report that raised them. It is the cheapest thing in this entire category to fix with software and one of the most valuable at inspection time.
Should you build custom or configure what you already own?
Some owners should not build, and we would say so before quoting. If you own one or two low hazard structures with a handful of instruments read monthly, a disciplined engineer and a well maintained spreadsheet is proportionate and honest. A custom system would add process without adding safety.
If your portfolio is fully automated, largely single manufacturer, and your regulatory reporting is light, look hard at what you can configure rather than build. Vista Data Vision does a genuinely good job of visualising and alarming on datalogger data and deserves consideration on that footing. Geokon and Campbell Scientific build excellent instruments and dataloggers, and their software is naturally organised around their own hardware, which works well when the portfolio matches. Worldsensing brings wireless sensing with a data layer on a similar pattern. None of them is trying to be the dam safety programme record, and for owners who do not need one, that is fine rather than a shortcoming.
Build when two or more of these are true. You hold more than roughly eight structures, or any structure with a high hazard classification and a population at risk. More than a third of your readings are manual and therefore live outside whatever automated system you have. Your thresholds sit inside datalogger programs and nobody can produce the rationale for them. An inspection has raised findings about data management or about closure of prior recommendations. You report to more than one regulator with different formats. Or your entire dam safety knowledge base is one engineer within five years of retirement.
How do hidden costs get into the quote?
Quotes in this category are usually wrong about the data and the field rather than about the application.
- Instrument diversity treated as a count. Several manufacturers and datalogger generations, plus instruments with no surviving documentation, means discovery per instrument type rather than per instrument.
- Historical migration priced as an upload. Importing a long record with correct datums, calibration history and declared data quality exclusions is careful engineering work and deserves its own line.
- Multiple regulatory regimes. A portfolio spanning hydro structures, state regulated dams and a tailings facility carries three different report shapes and three sets of expectations.
- Genuine offline capability. A field application that works with no signal for a full day is different from one that degrades, and the difference is real engineering.
- Survey and consultant ingestion. If reports arrive as documents, either agree a data format in the appointment or budget the extraction work.
A full platform adding reservoir correlated expected behaviour models, survey and inspection integration, regulator specific periodic report generation and portfolio risk views runs $150,000 to $350,000 phased across 6 to 12 months.
What separates a build that works from one that fails here?
Alerting that engineers do not mute. A fixed limit is the wrong instrument for this job, because a piezometer rising during a reservoir rise is normal and the same rise at steady pool after a dry month is exactly what the instrument exists to catch. Set the limit low and it alarms through every high pool event until nobody reads the emails. Set it high and it misses the signal. The version that earns its keep computes expected piezometric response given current and recent pool level and recent rainfall, compares it against measured, and triggers on the deviation. Ask a prospective developer how they would handle a piezometer rise during a filling event. If they separate expected from anomalous response and talk about correlating with pool level, they have done this. If they describe a threshold and an email, they will hand you alarm fatigue.
Escalation that reaches a person. An alert to a shared inbox is not a control. Route to a named engineer, require acknowledgement, escalate when it does not arrive, and keep the record, because an unacknowledged alert on a safety instrument is a finding in itself.
Availability that survives the event. Ask what happens when the network at a remote site is down and when the office system itself is unavailable, because a monitoring platform that is unreachable during the conditions it exists to watch has failed at its purpose.
And ownership settled in writing before kickoff: the repository, the database, the cloud accounts and an unrestricted export path. An instrumentation record is a safety document with a lifetime measured in decades and it has to outlive any vendor relationship, including ours. At Digital Heroes the client owns everything from the first commit, and a supplier who hesitates on that question is telling you something important about the next thirty years.
The evidence behind this guide
Independent findings on why this investment pays off. Every link goes to the primary source.
- SaaS spend averaged $4,830 per employee (up 21.9% year over year), with large enterprises (10,000+ employees) spending roughly $284M annually and running about 660 apps, while organizations wasted an average of $21M annually on unused licenses. Source: Zylo (2025) →
- This analysis cites IDC research that companies lose 20-30% of revenue annually to inefficiencies caused by data silos, Gartner's estimate that poor data quality costs organizations at least $12.9 million per year on average, and a Salesforce benchmark that 80% of IT leaders say data silos hinder digital transformation - illustrating the business case for integrating systems. Source: Cherry Bekaert (citing IDC, Gartner, Salesforce, DATAVERSITY) (2024) →
- Across more than 5,400 IT projects studied by McKinsey and the University of Oxford BT Centre, large IT projects ran on average 45% over budget and 7% over schedule while delivering 56% less value than predicted. Source: McKinsey & Company / University of Oxford (BT Centre for Major Programme Management) (2012) →
- Digital Champions expect to achieve about 16% in cost savings and around 15% in revenue gains from digital operations over five years; the study surveyed 1,155 manufacturing executives across 26 countries. Source: PwC / Strategy& (2018) →
Eleanor leads client services across the UK and EU, which means she sits between what a client asks for and what the delivery teams can realistically build. She writes about scoping, budget conversations and the questions worth asking before a build starts.
View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.
Frequently asked questions
Why do our piezometer alarms get ignored?
Because a fixed limit cannot distinguish a rise during a reservoir filling event, which is normal, from the same rise at steady pool after a dry month, which is the signal the instrument exists to catch. Set low it alarms through every high pool event until nobody reads it, set high it misses the event. Compute expected response from current and recent pool level and rainfall, compare against measured, and trigger on the deviation instead.
What can go wrong when we import forty years of readings?
Silent restatement. If the import applies today's calibration constants and today's surveyed reference to the whole series, every historical value shifts and the trend your engineer has relied on for a decade changes without anyone noticing, because the new numbers look plausible. Version instrument metadata with effective dates so past readings keep the conversion that applied when taken, and record which survey epoch each elevation came from.
How do we capture the readings taken by hand on a walk?
With an offline capable field application that stores locally, shows the previous value at the point of entry so an implausible number is caught on site, attaches photographs and syncs when signal returns. Manual readings are the majority at most portfolios, and an application that needs connectivity will be abandoned in week two. Manual and automated readings then get identical conversion, threshold and alerting treatment, which is the whole point.
What is the alert we are most likely to be missing?
The absence of readings rather than a threshold crossing. Telemetry links drop, modems fail and datalogger programs get edited during site visits, and nothing alerts because the last successful load looked normal. Give every instrument an expected reading frequency and raise an alert when readings stop arriving. Owners routinely discover multi week gaps weeks after the fact, and those gaps cover exactly the periods that later need explaining.
Why do inspections raise findings about our thresholds?
Because the values live inside datalogger programs, were set at commissioning by whoever configured the station, and nobody can produce the rationale. A threshold on a safety instrument is an engineering judgement and should carry the engineer of record who set it, the date, the reasoning and a revision history, held as governed data with controlled change. That also makes portfolio wide review possible, which is not achievable when values are scattered across station programs.
What usually goes wrong with prior recommendations at inspection?
Closure evidence. Recommendations are issued in a report, tracked in a spreadsheet or not at all, and the next inspection asks about every one of them. An owner who cannot show what was done about a recommendation issued five years ago has a much longer conversation than the underlying data warranted. Make recommendations tracked objects with owners, due dates, closure evidence and a link to the report that raised them.
Is Vista Data Vision or our instrument vendor's software enough?
For a fully automated, largely single manufacturer portfolio with light reporting, quite possibly, and it is worth configuring properly before commissioning anything. Vista Data Vision handles visualisation and alarming on datalogger data well, and Geokon, Campbell Scientific and Worldsensing organise naturally around their own hardware. None of them attempts the dam safety programme record: manual reading workflow, governed thresholds, recommendation tracking and regulator specific reporting.
What is usually underestimated in a quote for this?
Instrument diversity, which drives discovery per instrument type rather than per instrument, particularly where documentation no longer exists. Historical migration, which is careful engineering rather than an upload. Multiple regulatory regimes, since a portfolio spanning hydro, state regulated and tailings structures carries three report shapes. Genuine offline field capability. And survey or consultant data arriving as documents, which either needs a format agreed in the appointment or extraction work budgeted.
We run everything on spreadsheets and Airtable. How do we know it's time for custom software?
How much does a custom internal tool cost to build?
Can we start on Airtable or Retool now and move to custom software later?
How do I calculate whether custom software will pay for itself?
What does an internal tool cost for a small business with 20 to 50 employees?
Why do agencies charge for a discovery phase instead of quoting for free?
How do we migrate years of spreadsheet or Airtable data into a new internal tool?
Who can build a custom internal tools system?
Digital Heroes builds custom internal tools systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.
Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.
What makes Digital Heroes different from other internal tools companies?
Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.
Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.
How can I check Digital Heroes is legitimate before getting in touch?
Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.
Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.