Problems & solutions · Internal Tools

Continuity of Operations Software Problems: The 7 That Leave You Improvising, and How to Avoid Them

Continuity OF Operations Software workflow illustration showing common problems and fixes.
The short answer

The most expensive failure is a plan that has quietly stopped describing your organisation. At 2:40am, with file services encrypted, the only accessible copy is a printout from two years ago whose order of succession names a director who left in the spring and whose alternate facility for the assessor's office was repurposed as a records annex last year. Decisions then get made from memory in a group chat, the recovery takes longer than any objective in the document, and the audit finding afterwards is about currency, which is the one finding that costs you the next budget cycle as well as this incident.

Why does trying to model every department at once fail so often?

The biggest scope failure in continuity of operations (COOP) builds is starting with a complete organisational sweep. Forty departments, every essential function, every dependency, elicited in parallel. It sounds like the responsible approach and it is how these projects die, because each department is a discovery conversation with people who have day jobs, and the elicitation runs at the speed of the slowest calendar in the building.

Six months in, half the departments are documented, the first half is already out of date, and the sponsor is asking what was delivered. Nothing was, because a partially populated dependency model answers no question. The programme then gets a reputation for consuming time and producing nothing, which is much harder to recover from than a delayed release.

The alternative that works is narrow and provable. Take your three most critical essential functions and trace every dependency to a live system of record: the positions that perform them, the applications those positions use, the infrastructure and vendors underneath, the facility, the vital records. If any link in those three chains cannot be verified today, that is the project, and it is visible to a sponsor in weeks rather than quarters. A first release covering the function and dependency model, succession and delegation, vital records and drift detection runs $55,000 to $120,000 over 10 to 14 weeks in our delivery experience. Departments get added after the model has survived contact with three real functions.

What goes wrong when you migrate the existing plan into the system?

Every organisation wants to import the 200 page plan it already has, and the import is where the uncomfortable discoveries happen. A document can hold contradictions indefinitely because the contradicting facts sit ninety pages apart. A model cannot.

The classic case appears in the first week of migration. An essential function with a four hour recovery time objective depends on an application whose own objective is three days. Nobody wrote that deliberately. The function's objective came from asking a department head how quickly they needed to be back, and everyone answered immediately, so the plan carries forty first priorities with no arbitration between them. The application's objective came from the technology team, who were being realistic. Migration surfaces dozens of these at once, and if the project has not planned for the arbitration work, the team either stalls or, worse, quietly loads both numbers and the model inherits the contradiction it was built to expose.

The second migration trap is names. Plans are full of position titles that no longer exist, department names from a previous reorganisation, and people identified by first name and role. Matching those to the personnel system is manual, it is the highest value work in the whole migration, and it is routinely left to the end. Do it first. A position in the plan with no match in the human resources (HR) system is the single most important record in the build, because it is either a role that was abolished or a person who left, and both are live gaps.

Plan for arbitration sessions in the schedule, with a named executive who can rule on competing recovery objectives. Without that authority in the room, the model records the argument rather than resolving it.

Why do the personnel and asset inventory integrations break after launch?

The whole value of a continuity build is that the plan decays visibly rather than silently, and that depends on bindings to systems you already run. Those bindings are also where these projects break after go live.

The failure is rarely the interface. It is that the source system is not as authoritative as everyone assumed. The human resources system holds employment records but not the acting arrangements that matter during an event, so a plan bound to it shows a vacant position when in practice a deputy has been covering for eight months. The configuration management database is genuinely maintained for servers and genuinely fictional for business applications, so application dependency drift generates alerts that are all noise. The facilities system records space by building and room without any concept of a function being performed there.

What holds together is scoping the binding to the fields each system actually governs, and being explicit about the rest. Bind position holders to the human resources system. Bind acting and delegation arrangements to the continuity system itself, with an expiry, because no other system owns them. Bind infrastructure to configuration management and treat the business application inventory as continuity owned data until the technology team is ready to own it properly. Then send the continuity manager a weekly digest of ten specific drifts rather than an annual review cycle, because ten specific items get fixed and forty emails get eleven replies.

What happens when statutory succession, delegation and vital records are not covered?

Public sector, utility and health system continuity carries constructs that corporate business continuity models treat as optional. Orders of succession set by ordinance. Delegations of authority with specific triggers and limits. Devolution to a geographically separate site. Vital records defined by a records retention schedule. Configured in as custom fields, they become documentation. Modelled properly, they become operational.

Delegation is the one that hurts during a real activation. Most plans describe it in a paragraph. The live question is sharp: who can authorise emergency procurement above the normal threshold, who can direct staff to relocate, who can declare the event, and what happens when that person is unreachable. Model it as a rule with a trigger, a scope, a limit and an expiry, then record who actually exercised what and when. That record answers the auditor and the elected official asking who approved a particular expenditure at 4am, and during the event it prevents the paralysis of nobody being sure they are allowed to act.

Vital records are the half nobody funds. Essential functions get attention because they are visible. The deeds, case files, licences, personnel records and utility as built drawings get a table, and the function cannot be performed without them. Each needs a format, a storage location, a copy that is not in the same building or the same cloud tenant, a restoration procedure someone has actually run, and an owner who is a person rather than a department name. Tie them to the retention schedule you already maintain rather than building a second list, because a second list diverges within a year. Then sample test restorations quarterly and record the result, since an untested backup is a belief rather than a control.

Should you build custom or configure what you already own?

Fusion Risk Management and Castellan are capable business continuity platforms with genuine dependency modelling, and for a mid sized private organisation they are the right purchase. Riskonnect approaches continuity as one module within governance and risk, which suits an organisation buying the whole suite. Veoci is flexible and many emergency managers already run other functions on it, so continuity becomes configuration inside an environment their team knows.

Configure rather than build if you already own one of these and your continuity requirement is corporate in shape. Configure rather than build if you are small enough that one person genuinely holds the picture, if your plan is under fifty pages, or if your regulatory driver is a checkbox rather than a real audit. Software will not manufacture discipline you do not currently have, and a maintained document plus an honest annual tabletop is proportionate for a single agency with fifteen staff and one alternate site.

The reasons buyers move away from configuration are two, and both are structural. Continuity touches every department, so per user pricing across a forty department county or a multi campus health system becomes a large recurring line for software most of those users open twice a year. And the statutory constructs above are native to your obligations and configured into theirs. Build when the plan spans dozens of departments, when those constructs must be modelled rather than described, when the plan needs to reflect systems that change weekly, when you have failed an audit on currency, or when a real activation has already shown you the document was unusable.

How do hidden costs get into the quote?

Continuity quotes are unusually prone to underestimation because most of the work is not software.

  • Department elicitation counted as one line. Each department is a discovery conversation, a set of essential functions and an argument about recovery objectives. Department count is the primary cost variable, not feature count.
  • Inventory quality assumed. An organisation with a maintained configuration management database has an integration. An organisation whose application inventory is a spreadsheet last touched under a previous administration has an inventory project inside the continuity project, and it is better to admit that in week one than discover it in week six.
  • The offline export treated as a report. An automatically generated, current, offline usable version of every plan is a core requirement, and it has to be tested the way a backup restoration is tested.
  • Exercise management deferred. Exercises produce the only real data the programme has, and bolting the module on later means the first year of exercises is not captured.
  • Recovery objective arbitration. Somebody senior has to rule between competing priorities, and that time belongs in the plan.

The full platform adding exercise management with corrective action tracking, activation mode with role based task assignment, devolution and alternate facility planning, departmental self service and audit reporting runs $140,000 to $320,000 across 6 to 11 months.

What separates a build that works from one that fails here?

Three things, and none of them is a feature list.

The first is that the system detects contradictions rather than storing them. Ask a prospective developer to model, on a whiteboard, an essential function whose recovery objective is shorter than one of its dependencies. If they show you how the system raises and surfaces that conflict the moment it is created, along with positions that have no named successor, applications with no owner and vital records with no verified separate copy, they understand the assignment. If they show you a form for entering functions, you are commissioning a document with a database behind it.

The second is that the system works when it is not available. A continuity platform that depends on the infrastructure it is meant to help recover has failed at its only job. There must be an exportable, current, offline usable version of every plan, generated automatically and held somewhere independent, and it should be verified on a schedule.

The third is that exercises feed the objectives. Capture what was tested, who took part, what actually failed, the recovery times observed rather than planned, and corrective actions with owners and due dates. Then push observed times back into the model, because a function that has never recovered faster than eleven hours across three exercises does not have a four hour objective, it has an aspiration and a documented history of missing it. Auditors want evidence of exercise and closure, not a well written plan.

Settle ownership before kickoff, including exported plans and exercise history. At Digital Heroes the organisation owns the code and the data from the first commit.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. Technology 'Leaders' grow revenue at more than twice the rate of 'Laggards'; laggards surrendered 15% in foregone annual revenue in 2018 and stood to miss out on as much as 46% in revenue gains by 2023 if they did not change their enterprise technology approach. Based on a survey of more than 8,300 organizations across 20 industries and 20 countries. Source: Accenture (2019) →
  2. SaaS spend averaged $4,830 per employee (up 21.9% year over year), with large enterprises (10,000+ employees) spending roughly $284M annually and running about 660 apps, while organizations wasted an average of $21M annually on unused licenses. Source: Zylo (2025) →
  3. Only 22% of firms are 'future ready' having significantly transformed digitally; these companies show average revenue growth 17.3 percentage points and net margins 14.0 percentage points above their industry average. Source: MIT Center for Information Systems Research (MIT Sloan) (2022) →
  4. The right combination of digital transformation actions can unlock as much as US$1.25 trillion in additional market capitalization across Fortune 500 companies, while the wrong combinations put more than US$1.5 trillion at risk; companies with all three core factors (strategy, aligned technology, and change capability) saw a 5% market-value lift relative to peers. Source: Deloitte (2023) →
Oliver H. · Senior Account Director · UK · London

Oliver runs UK client accounts day to day, chairing the calls where scope, budget and timeline meet reality. He is useful reading for anyone about to commission custom software and wondering what a healthy agency relationship should feel like from the client side.

View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.

FAQ

Frequently asked questions

Why does our COOP plan stop being accurate within months of the annual review?

Because a document has no way to notice the world moved. Directors leave, applications are migrated, facilities are reassigned and vendor contracts lapse, and none of it reaches a Word file until someone runs a review and sends emails. Binding the plan to live personnel, asset and facility systems turns silent decay into a weekly digest of ten specific drifts, which a continuity manager will actually work through, unlike forty emails that produce eleven replies.

What surfaces first when we migrate our existing plan into a model?

Contradictions the document was hiding, usually a function with a four hour recovery objective that depends on an application whose own objective is three days. Those appear in bulk because department heads all answered immediately when asked how quickly they needed to be back, so the plan carries dozens of first priorities. Schedule arbitration sessions with a named executive who can rule between competing objectives, or the model will record the argument instead of resolving it.

Which departments should the first release cover?

None of them, in the sense of a sweep. Take your three most critical essential functions and trace every dependency to a live system of record: positions, applications, infrastructure, vendors, facility and vital records. Any link in those chains that cannot be verified today is the project. Attempting forty departments in parallel means a partially populated model that answers nothing, six months of elicitation, and a sponsor asking what was delivered.

Why do the human resources and asset integrations produce noise?

Because the source systems are less authoritative than everyone assumed. The personnel system holds employment records but not the acting arrangements that matter during an event. The configuration database is genuine for servers and largely fictional for business applications. Bind only the fields each system actually governs, hold acting and delegation arrangements in the continuity system itself with expiry dates, and own the application inventory yourself until the technology team is ready to.

How should delegation of authority be modelled rather than described?

As a rule with a trigger, a scope, a limit and an expiry, plus an activation record of who exercised which authority and when. During a live event the questions are immediate: who can authorise emergency procurement above the normal threshold, who can direct staff to relocate, who can declare the event, and what happens when that person is unreachable. Afterwards, the same record is what the auditor and the elected official ask for.

Is Fusion or Castellan enough for a county or health system?

They are capable platforms with real dependency modelling and a sound purchase for a mid sized private organisation. Public sector and health system buyers hit two structural walls: per user pricing across dozens of departments for software most of those users open twice a year, and a corporate model where orders of succession, delegations of authority and vital records defined by a retention schedule are configured in rather than native to the design.

What costs are usually missing from a continuity software quote?

Department elicitation, which is the primary cost variable rather than feature count. Inventory remediation, when the application list is a spreadsheet from a previous administration, which is a project inside the project. The automatically generated offline export, which is a core requirement and needs testing like a backup restoration. Exercise management, which is often deferred so the first year of exercises goes uncaptured. And the executive time to arbitrate competing recovery objectives.

What should we ask a developer to prove they have done this before?

Ask them to model, on a whiteboard, an essential function whose recovery objective is shorter than one of its dependencies. The right answer shows how the system raises that conflict the moment it is created, alongside positions with no successor, applications with no owner and vital records with no verified separate copy. A form for entering functions means you are commissioning a document with a database behind it, which is what you already have.

How much does a custom internal tool cost to build?
Most custom internal tools cost $8,000 to $40,000 to build, based on Digital Heroes delivery data across 2,000+ client projects. A single-purpose tool like an approval dashboard or inventory tracker sits at the low end, while a multi-department platform with role-based access and several integrations pushes past $40,000. The three biggest cost drivers are the number of user roles, the number of systems the tool must connect to, and custom reporting requirements.
Should we build our internal tool in Retool instead of hiring developers?
Retool is the right choice if someone on your team is comfortable with SQL and JavaScript and the audience is a handful of technical users, because a basic CRUD dashboard comes together in days. Hire developers when non-technical staff will use the tool daily, when the logic goes beyond forms sitting on a database, or when per-seat pricing stings, since Retool's Business tier lists at $50 per standard user per month. A pattern Digital Heroes sees often: companies arrive after a year on Retool with a tool nobody can maintain because the one person who built it has left.
How long does it take to build an internal tool from scratch?
A working first version typically ships in 4 to 8 weeks, and larger multi-module tools run 10 to 16 weeks. Across Digital Heroes internal tool projects the schedule splits into roughly one week of process mapping, 3 to 6 weeks of build, and 1 to 2 weeks of testing with your actual staff. The most common delay is not development but waiting on the client for sample data and workflow decisions, so name one internal owner before kickoff.
What tech stack should an internal tool be built with?
Boring and popular: a React or Next.js frontend, a Node.js or Python backend, and PostgreSQL covers the vast majority of internal tools and keeps future hiring easy. The stack matters far less than whether a different developer can pick the code up in two years, so require documentation as a deliverable and avoid anything exotic. Treat it as a red flag if an agency pushes a proprietary platform only they maintain, because that quietly converts your tool into a subscription to that agency.
When does a company outgrow Airtable?
The usual breaking points are record limits, permissions, and automation complexity. Airtable's Team plan caps each base at 50,000 records and Business at 125,000, so operations logging thousands of rows a month hit the ceiling within a year or two. The other trigger Digital Heroes sees constantly is permissions: restricting who can view specific fields or records is clumsy below Airtable's Enterprise tier, which becomes a genuine problem once salaries, pricing, or client contracts live in the base.
Who owns the code when an agency builds our internal tool?
You should, outright, with full IP transfer in the contract and the code delivered to a repository you control, such as your own GitHub organization. Digital Heroes transfers complete ownership on final payment as standard practice, and any agency that keeps the code or licenses it back to you is building a dependency you will pay for later. Confirm you also own the hosting, domain, and database accounts, since many of the vendor disputes Digital Heroes gets called into involve infrastructure registered under the agency's name.
How do I know when spreadsheets are no longer enough to run my operations?
Replace the spreadsheet once more than three people edit it, versions travel by email, or a single broken formula could cost real money. Other reliable signals: staff keep personal shadow copies, month-end reporting takes days of manual assembly, and nobody can say who changed a number or why. In Digital Heroes discovery calls the tipping point is almost always a specific expensive error, a mispriced quote, a missed order, or payroll built on a tab someone sorted wrong.
What are the biggest mistakes first-time software buyers make?
Choosing the lowest bid, paying more than 30-40% upfront instead of on milestones, skipping a written specification, and having no maintenance plan for after launch. The most expensive of the four in Digital Heroes rescue projects is the missing spec: without written acceptance criteria, done becomes an argument instead of a checklist, and every disagreement resolves in the vendor's favor. Fix those four and you have avoided most of the ways these projects fail.
How many developers does it take to build an internal tool?
Two to four people covers nearly every internal tool: one or two developers, a part-time designer, and a project manager who doubles as your single point of contact. Internal tools rarely need consumer-product polish, so a full-time dedicated designer is usually wasted budget. On Digital Heroes projects, a two-person core team handles the typical 4 to 8 week build, with a specialist pulled in briefly for a tricky integration or a security review.
How do I vet a development agency for an internal tools project?
Ask to see two or three internal tools they have shipped and whether those clients still use them daily, because internal tools fail on adoption, not code quality. Good signs: they ask to see your current spreadsheet or process before quoting, they propose a phased build instead of one big launch, and they spell out who handles training and post-launch changes. Walk away from anyone who gives a fixed price before seeing your actual workflow, since internal tools live or die on process details.
Does it matter which tech stack the agency wants to use?
Yes, but not in the way most buyers expect: the goal is boring, popular technology such as React, Node.js or Python, and PostgreSQL, because any future team can maintain it and hiring a replacement developer takes days, not months. The red flag is an agency-proprietary framework or an unusual language, which welds you to that one vendor no matter what your contract says about code ownership. A useful test: could you find three freelancers fluent in this stack within a week? If not, push back.
Can a custom internal tool connect to QuickBooks, Salesforce, and the other software we already use?
Yes, and integrations are usually the strongest argument for going custom instead of chaining tools together with Zapier. QuickBooks, Salesforce, Shopify, Stripe, Slack, and Google Workspace all have mature APIs, and each integration typically adds $1,500 to $5,000 to a Digital Heroes build depending on how much two-way syncing you need. The honest caveat is legacy industry software without an API, which may need file-based imports instead of a live connection, so list every system in the first conversation.
Who can build a custom internal tools system?

Digital Heroes builds custom internal tools systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other internal tools companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading
let's build

Build something worth launching.

A plan, a team, a timeline, within 24 hours. No decks, no discovery calls. Tell us what you're building and we'll come back with a real scope and a real number.

message us directly · we reply within one business day

mission briefing

Monthly dispatch

Playbooks, real build costs, and what we're shipping. One email a month. No fluff.

visit us

New York HQ

1140 Broadway, Suite 704 · New York, NY 10001

Get directions
Online now

Hey there 👋 How can we help you today?