Problems & solutions · Field Service Management

Telecom Site Power and Generator Monitoring Software Problems: The 7 That Cost Real Money, and How to Avoid Them

Telecom Site Power Monitoring Software software overview illustration showing common problems and fixes.
The short answer

The most expensive failure is a system that raises alarms without knowing what each site is worth. A mains failure on a site with a working generator, a full tank and a healthy string, and a mains failure on a battery only site whose string has lost half its capacity and which backhauls a hospital's fibre, arrive on the network operations wall looking identical. Once that happens at scale the operations team starts filtering power alarms, which is the real failure state: not a missing alarm, an ignored one. Then a storm line crosses the region at two in the morning, sixty three power alarms appear, the on call technician is dispatched to whichever site is closest to his house, and by five the site that actually mattered has been down for two hours and you are drafting an outage narrative for a tenant whose contract carries service credits.

Why does the protocol integration scope failure happen so often?

Because the interesting part of the brief is analytics and the expensive part is talking to the estate. Battery health scoring and fuel modelling demonstrate well in a proposal. Getting real data out of a fleet assembled over fifteen years of build programmes and acquisitions does not, so it gets priced as "device integration" and the number is wrong.

What is actually there is a serial link into a rectifier controller with a vendor specific register map, a newer controller speaking a network management protocol, an inherited site from a utility partnership speaking something else again, a generator controller speaking its own protocol over a cellular modem, and a handful of sites where the only reliable data path is a technician's phone. Each controller family is genuine integration work, typically one to three weeks including a lab unit and a site validation, and the lab unit is not optional because a vendor's documented register map and the map implemented in a specific firmware version are not always the same document.

The fix is to scope the first release around the two or three highest count families, which usually covers most of the fleet, and to fund the long tail as backlog rather than as a launch dependency. Ask a candidate developer which controllers they have driven on a live site, by vendor, model and transport. A network connection to a modern controller and a serial link through a converter at a site with marginal signal are different projects, and only one of them teaches anything useful.

What goes wrong when you migrate site asset and battery records?

The site list is easy and misleading. What the analytics actually depend on is the battery record, and in most fleets that record is worse than the site record by a wide margin.

Strings get replaced by contractors who record the work in a job ticket rather than in an asset register, so install dates are wrong for exactly the strings that were replaced most recently. Rated capacity is often the original design figure rather than what is physically installed after a partial swap. Mixed age strings sit in the same cabinet because two of four blocks were changed under warranty. Sites that were rebuilt during a technology upgrade carry the old configuration in the register and a different one in reality.

The consequence is specific. Any state of health method compares measured behaviour against rated capacity, so a wrong rated capacity produces a wrong score with full confidence, and your replacement list gets ordered by a number that is quietly fiction.

The fix is to treat the battery as an asset with its own install, replacement and test history rather than as an attribute of the site, and to fund a register reconciliation as part of the first release. Then design the scoring so that a site with no credible baseline is flagged as unknown rather than scored, because an honest gap in the list is far more useful to a capital planner than a confident wrong answer.

Why do the integrations that matter here break after launch?

Because the estate is outdoors and the events you care about are exactly the events that break connectivity.

Backhaul is the first. A site loses its link precisely when it goes to battery in a storm, which means a system without local buffering loses the discharge curve for every outage that mattered and keeps the curves for the trivial ones. The symptom is a dataset that looks complete and is systematically biased toward benign events. Ask any prospective developer about store and forward first. If it is not in the answer, they have built for data centres.

Field dispatch is the second. Predicted time to site down is only useful if it reaches whoever schedules technicians, which means an integration into your field service scheduling and a criticality tiering that today lives in someone's head. Both of those are your business rather than a vendor's, and both tend to be assumed rather than scoped.

Fuel supplier dockets are the third. Delivery records arrive as documents, sometimes as photographs, on the supplier's timetable rather than yours, and reconciliation depends on them arriving reliably. When a supplier changes their paperwork or their delivery contractor, the reconciliation quietly stops finding exceptions and looks like good news.

The fix is monitoring on the inputs, not just on the sites. A collector that has gone quiet, a docket feed that has produced nothing for a fortnight, and a controller family whose polling success rate has dropped are all operational alerts in their own right.

What happens when criticality tiering and fuel reconciliation are not covered?

You keep the alarm problem you started with, and you keep overpaying for diesel.

Criticality first. Without a tier that reflects your tenants, your backhaul topology and any public safety or emergency service obligations, every alarm has equal weight, and a hub site carrying twelve downstream sites looks identical to a leaf. The consequence is not that alarms are wrong, it is that they become noise and get filtered, after which the monitoring system is decorative. Writing down criticality tiering is unglamorous work that only your organisation can do, and it is worth doing whether you buy or build.

Fuel second. Scheduled delivery produces two failures at once: sites in stable grid areas get topped up while still nearly full, and sites in load shedding or storm exposed areas run dry during the week they were needed. Meanwhile the gap between fuel purchased and fuel that could plausibly have been burned given actual run hours is where losses live, and on a fleet with remote sites and a contracted delivery arrangement nobody reconciles it.

The fix on criticality is to fuse telemetry with the facts only you hold, and to output predicted time to site down rather than a binary alarm, suppressing mains failures on sites with autonomy above a threshold until the threshold is crossed. The fix on fuel is a consumption model per site derived from its own history, litres per run hour at its typical load, reconciled against delivered litres on the docket and against the tank level step the delivery should have produced. A docket claiming four hundred litres into a tank that rose by two hundred and twenty becomes an exception before the invoice is paid.

Should you build custom or configure what you already own?

A large share of operators should configure, and the threshold is reasonably clear.

If you run under roughly 150 sites on a substantially single vendor rectifier estate, use that vendor's own monitoring. Vertiv Environet, Schneider EcoStruxure and Eaton Brightlayer present the state their own hardware reports properly, the integration is already done, and a custom build would be an expensive route to a similar screen. Spend the difference on batteries and on tank senders at the sites where fuel volume justifies them. The same applies if your sites are grid stable and mostly battery only with no generators, because two of the three problems above simply do not apply to you.

Configuration you have probably not done: setting alarm thresholds per site class rather than globally, recording criticality tiering anywhere at all, and turning contractor job tickets into a battery asset register. Those three will improve dispatch quality inside the tools you already own.

Build when three things are true together. You have a genuinely multi vendor fleet, so no single platform sees the whole estate. Your fuel and battery spend is large enough that a modest percentage improvement is a real number. And you carry tenants or obligations where a site down event has a contractual cost. Tower companies and neutral host operators hit all three early because their commercial model is uptime sold to multiple tenants with different criticality. Mobile operators hit it once the operations team has started filtering power alarms, which is a symptom you can go and look for this week.

How do hidden costs get into the quote?

Five places, and the first two are the ones that move the number most.

Controller families and firmware vintages. Each family is one to three weeks with a lab unit and a site validation, so a quote that says "multi vendor support" without a named list has priced the two easy ones. Ask for the list and the price per family.

Second, hardware that has to be fitted before a site can report anything. Tank senders, string level monitoring where only a shunt exists, and cellular connectivity at sites that currently have none are capital items outside the software quote, and they determine which of the promised features actually work on which sites.

Third, tenant service level evidence. If you carry contractual service credits, the reporting has to be defensible rather than indicative, which means immutable event records and a documented derivation, and that is more work than a dashboard.

Fourth, multi country deployments, where fuel procurement, grid behaviour and load shedding patterns differ enough to need separate models rather than one with a region field.

Fifth, the ongoing line. Firmware changes, sites get rebuilt, and suppliers change their paperwork. A quote with no maintenance figure has moved that cost rather than removed it.

What separates a build that works from one that fails here?

One output, delivered early. In Digital Heroes delivery experience the builds that work start with one region, the top two hardware families and the single question of predicted time to site down, and ship edge collection with store and forward, battery health from opportunistic discharge, fuel modelling with delivery reconciliation and criticality ranked triage in 12 to 18 weeks. That one output changes dispatch behaviour on day one. Builds that fail start with a fleet wide dashboard and reach month six with beautiful visibility into an estate they still cannot act on.

The second differentiator is whether the system treats every unplanned mains failure as a free capacity test. Capturing the discharge curve, the load in amps, the ambient temperature and the time to recovery gives you a capacity trend per string over a year without sending anyone anywhere, and it flags the sites that never discharge as the small minority worth a deliberate test. That is what turns a blanket age based replacement programme into a risk ranked list.

The third is a vetting question. Ask a developer to describe the state of health method they would use before they have seen your data. A credible answer talks about temperature compensation, load at the time of discharge, and trending per string. An answer that is a voltage threshold is a repackaged alarm.

Last, settle ownership of the code, the cloud accounts and the device credentials before kickoff. This system accumulates years of discharge and burn history that becomes the evidence base for your battery and generator capital decisions, and that history should never sit inside a vendor tenancy.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. PTC identifies the leading causes of failed first visits as parts unavailability (the single most-cited complaint, named by 51% of field service executives), technicians lacking the required equipment or skills, and insufficient time allocated to the job - making parts logistics and skills-based dispatch the highest-leverage fixes. Source: PTC (2023) →
  2. Salesforce's field-service research (State of Service / field service trends, survey of 5,500+ service professionals) found that 74% of mobile workers report increasing workloads and 47% say appointments don't go as planned due to customer miscommunication, unaccounted-for parts, or insufficient appointment lengths and travel times. (The separate claim that admin tasks consume ~30% of a technician's hours is NOT supported by the report - the seventh-edition data instead states technicians spend about 18% of working hours, ~7 hours/week, on admin, and only ~32% of time interacting with customers.). Source: Salesforce (2024) →
  3. 76% of developers are using or planning to use AI tools in their development process in 2024 (up from 70% in 2023), with current active use rising to 62% from 44%; 81% agree increasing productivity is the biggest benefit of AI tools. Source: Stack Overflow (2024) →
  4. Across more than 5,400 IT projects studied by McKinsey and the University of Oxford BT Centre, large IT projects ran on average 45% over budget and 7% over schedule while delivering 56% less value than predicted. Source: McKinsey & Company / University of Oxford (BT Centre for Major Programme Management) (2012) →
Vikram R. · VP Engineering · Delhi

Vikram runs the engineering function at Digital Heroes, from how teams are structured to how code gets reviewed and released. He writes about the trade offs behind build decisions: what to buy, what to build, and where technical debt is worth taking on deliberately.

View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.

FAQ

Frequently asked questions

Our operations team filters power alarms. Is that a process problem or a software problem?
It is a software problem with a process symptom, and it is the clearest signal that you have outgrown asset state monitoring. Alarms get filtered when a mains failure on a healthy site with a full generator tank looks identical to one on a battery only site with a degraded string. The fix is to output predicted time to site down rather than a binary alarm, ranked against site criticality, and to suppress events on sites with autonomy above a threshold until that threshold is crossed.
Can we assess battery health without taking live sites to battery deliberately?
Yes, and it is the highest value capability in the category. Treat every unplanned mains failure as a free capacity test by capturing the discharge curve, load in amps, ambient temperature and time to recovery, then trend it per string against rated capacity. Sites that never discharge get flagged for a deliberate test precisely because they generate no data, which is a small minority worth a crew visit. That converts a blanket age based replacement programme into a risk ranked list.
Why would a state of health score be wrong even when the telemetry is good?
Because the score is measured against rated capacity, and most fleets have a battery register that does not reflect what is physically installed. Contractors record replacements in job tickets rather than asset registers, partial swaps leave mixed age strings in one cabinet, and rebuilt sites carry the old configuration on paper. Model the battery as an asset with its own install and test history, and design the scoring to report unknown rather than guess where no credible baseline exists.
What happens to our data when a site loses backhaul during an outage?
Without local buffering you lose the discharge curve for exactly the events you care about, and you keep the curves for the benign ones, which produces a dataset that looks complete and is systematically biased. The edge collector must store locally and replay on reconnect. Ask any prospective developer this before anything else, because if store and forward is not already in their answer they have built for data centre halls rather than outdoor sites.
How do we catch fuel losses across remote generator sites?
Reconcile three numbers that should agree: predicted burn from a consumption model derived from that site's own history, delivered litres on the docket, and the tank level step the delivery should have produced. A docket claiming four hundred litres into a tank that rose by two hundred and twenty becomes an exception before the invoice is paid, and a site burning fuel with no corresponding run hours is either a leak or a siphon. This needs a tank sender at the sites where fuel volume justifies the hardware.
How long does controller integration take per hardware family?
Plan one to three weeks per family including a lab unit and a live site validation, and do not skip the lab unit, because a vendor's documented register map and the map implemented in a given firmware version are not always the same document. Scope the first release around the two or three families covering most of your fleet and fund the long tail as backlog. A quote offering multi vendor support without a named list of controllers has priced the easy ones.
Should this drive dispatch or just raise alarms?
It should drive dispatch, and that is where it stops being a dashboard. The useful output is predicted time to site down, computed from current load, battery state of health and generator status, ranked against site criticality and travel time from where your technicians actually are. That requires an integration into whatever scheduling system your field team uses and it requires criticality tiering to be written down, which is work only your organisation can do and is worth doing either way.
We have 120 sites on one rectifier brand. Do we need any of this?
Probably not. On a coherent single vendor estate at that size, the vendor's own monitoring plus a disciplined battery replacement programme is the right spend, and the money is better used on strings and tank senders. The picture changes when you cross into multiple hardware families, when generators and fuel become a material cost line, or when you carry tenants whose contracts include service credits for downtime. Two of those three usually arrive together.
Can we migrate years of data out of our current system into new custom software?
Almost always yes, through CSV exports or the vendor's API, and migration should be scoped as its own workstream with field mapping, a dry run, and a planned cutover window rather than an afterthought. The real time sink is rarely moving the data; it is cleaning it, since years of duplicates, free-text fields, and inconsistent formats surface all at once. Pull a full export from your current vendor before committing to anything new, because some SaaS plans restrict exports on lower tiers.
Can custom software connect to the tools we already use, like QuickBooks, Stripe, and Google Workspace?
Yes, and connecting your existing tools is one of the main reasons to build custom: mainstream platforms like QuickBooks, Stripe, Shopify, and Google Workspace all publish documented APIs. Budget 1 to 3 weeks of work per integration depending on API quality and how much data flows in both directions. Ask any vendor whether they have integrated with your specific tools before, because quirks like QuickBooks' OAuth token handling and API rate limits get learned on someone's project, and it should not be yours.
Do my field technicians need a native mobile app, or will a web app work?
If your technicians ever work in weak signal, you need a native or offline-capable app, because a plain web app fails exactly where field work happens: basements, mechanical rooms, and rural routes. Cross-platform frameworks like React Native or Flutter give one codebase for iPhone and Android with full offline storage, which is how Digital Heroes builds most technician apps. A web app is the right call for the office dispatch console, where connectivity is guaranteed.
Should we start with an MVP or build the full field service platform in one go?
Start with an MVP that can run one real crew for one real week: scheduling, dispatch, job completion with photos and signatures, and invoicing. That slice typically costs $40,000 to $70,000 and ships in about 12 weeks, and technician feedback then decides phase two. Teams that built the full platform up front reworked 30 to 40 percent of it after field use in Digital Heroes experience, which is the most expensive way to discover what dispatchers actually need.
How much would it cost to build something like ServiceTitan just for my company?
A true ServiceTitan clone would cost millions and you do not need one, because companies that bring this request to Digital Heroes typically use 20 to 30 percent of its features. Building that slice, shaped to your exact dispatch board and technician day, runs $80,000 to $200,000 depending on offline requirements and integrations. The field service builds that succeed copy a workflow, not a product.
Is custom software more secure than off-the-shelf SaaS?
Neither is secure by default; security tracks the practices of whoever builds and operates the system, not the model. SaaS gives you the vendor's certifications and patching but puts your data in a shared multi-tenant platform on their terms, while custom gives you full control over data residency, access rules, and compliance requirements like HIPAA, with the responsibility sitting with you and your agency. Before hiring anyone for a system holding sensitive data, ask for their security checklist: encryption at rest and in transit, an OWASP Top 10 review, role-based access, and a penetration test before launch.
Should I hire a freelancer or an agency to build my field service software?
An agency in almost every case, because a field service build spans a mobile app, a dispatch web console, a backend, offline sync, and accounting integrations, which is four or five specialties one person rarely covers. A freelancer is the right choice for a single integration or a well-scoped add-on under $15,000. The solo-built field service systems Digital Heroes inherits fail most often at handover, when the freelancer has moved on and nobody can safely modify the sync engine.
How many people should be working on my software project?
Three to five for a typical focused build: a project lead, one or two engineers, a designer, and part-time QA, which is the standard shape across 2,000+ Digital Heroes projects. Larger platforms justify 6 to 10, but a ten-person team on a small first version usually signals bill padding rather than horsepower. What predicts success is whether a senior engineer is writing your code daily, not the headcount on the proposal.
How does custom field service software work when technicians have no cell signal?
Properly built field software stores the technician's entire day on the device, including job details, forms, photos, signatures, and parts, then syncs automatically when signal returns. The hard engineering is conflict resolution: deciding what happens when a dispatcher reassigns a job while the technician is working it offline. That logic has to be designed before the build starts, because retrofitting offline into an app that assumed a connection is close to a rewrite.
Who can build a custom field service management software system?

Digital Heroes builds custom field service management software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other field service management software companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading
let's build

Build something worth launching.

A plan, a team, a timeline, within 24 hours. No decks, no discovery calls. Tell us what you're building and we'll come back with a real scope and a real number.

message us directly · we reply within one business day

mission briefing

Monthly dispatch

Playbooks, real build costs, and what we're shipping. One email a month. No fluff.

visit us

New York HQ

1140 Broadway, Suite 704 · New York, NY 10001

Get directions
Online now

Hey there 👋 How can we help you today?