Industry guide · Internal Tools

Communications Surveillance and Archiving Software: How Do You Produce Every Message in a Defensible Form When the Request Arrives?

Communications Surveillance Archiving software visual showing messages square, text search, and archive.
The short answer

If you are a broker dealer, bank or investment adviser capturing more than about five channels across several jurisdictions, and your supervisory review is a spreadsheet of sampled messages signed off by a principal each month, build the review and case layer. A focused first release covering unified message ingestion from your existing capture connectors, a policy driven review queue with documented sign off and defensible production runs $95,000 to $210,000 and ships in 14 to 20 weeks in our delivery experience. A full platform adding voice transcription review, risk scoring across channels, legal hold, jurisdictional retention rules and trading linkage runs $260,000 to $700,000 phased over 9 to 16 months. If you are a single jurisdiction firm on email and one chat platform, Smarsh or Global Relay end to end is cheaper and entirely adequate.

Why capture is solved and supervision is not

The request arrives on a Tuesday. All communications involving four named individuals across an eleven month window, any channel, produced in original form with metadata intact. Your archive holds email properly. Chat is in a different vendor. Voice is in the turret recorder with a separate retention schedule. Teams messages started being captured in March and the two months before that are gone. One of the four people was a contractor whose messages were captured under a different identity. Producing this is now a six week project involving three vendors and a law firm, and the gaps will be found by someone other than you.

Enforcement over off channel communications has been the clearest compliance signal of recent years, with actions across dozens of firms and penalties large enough that capture programmes became board level obligations rather than compliance projects. The industry responded by buying capture, which was the right first move. Smarsh, Global Relay, Theta Lake, Shield and Behavox all capture well, and rebuilding connectors to WhatsApp, Bloomberg chat, Microsoft Teams, Zoom and a voice turret would be an expensive way to reinvent something mature.

Capture, however, is the easy half. What regulators actually examine is supervision: whether a qualified person reviewed the right things, on a defined basis, with documented reasoning, and whether the firm can prove it. That part is where packaged tools hand you a generic lexicon, a queue and a checkbox, and where firms discover that their review programme is a monthly ritual nobody believes in.

Problem 1: lexicon review generates volume, not supervision

A keyword list flags every message containing guarantee, and most of them are somebody guaranteeing to call back after lunch. Reviewers learn the pattern within a week and start closing in bulk. The queue is worked, the sign off is recorded, and the programme has quietly become a formality.

What changes it is treating review population as a policy decision rather than a search result. A build lets compliance define populations explicitly: everything from this desk during the quiet period, all external messages from staff on the restricted list, a random sample of a stated percentage from every registered person, plus risk triggered items. Each population has a rationale, a reviewer role, a frequency and an evidentiary record. When an examiner asks how you decided what to review, you show the policy and its version history rather than describing a keyword list.

This is where language models earn their place, and it is the one part of communications compliance where they clearly outperform the incumbent approach. Keyword matching cannot tell frustration from a complaint, or a joke about a market from a discussion of a specific order. A classifier scoring messages for the behaviours your policy actually cares about, running alongside the lexicon rather than replacing it, reduces reviewed volume substantially while raising the proportion of reviewed items that are genuinely worth reading. It must remain advisory: the model prioritises, the qualified principal decides and signs.

Problem 2: identity is fragmented and that is where the gaps live

The same person is an email address, a Bloomberg UUID, a Teams object identifier, a mobile number, a WhatsApp account and a turret extension. Some of those changed when they married, moved desks or came back as a contractor. Reconstructing a person's full communication history therefore depends on an identity map that in most firms is a spreadsheet maintained by whoever remembers.

A build makes identity a governed record: every channel identifier bound to a person with effective dates, sourced from the human resources (HR) system and the registration records so joiners and leavers propagate automatically. That single structure is what makes a production request answerable in hours instead of weeks, and it is also what makes supervision honest, because a reviewer covering a desk actually covers everyone on it including the person who joined last Thursday. The failure mode is always the same: the message existed, the capture worked, and nobody knew that identifier belonged to that person.

Problem 3: retention and legal hold rules collide across jurisdictions

Books and records obligations under SEC Rule 17a-4 sit alongside adviser record keeping requirements, market abuse obligations in the EU and UK, and privacy regimes that push in the opposite direction. A message from a trader in Frankfurt to a client in Singapore, held on a US archive, is subject to several regimes at once, and the answer to how long you keep it and who may read it is not uniform. Meanwhile a legal hold must freeze deletion for a specific matter without breaking the general schedule for everything else.

Packaged archives implement a retention policy. What they handle less gracefully is a matrix of overlapping policies where the longest applicable period wins, holds override schedules, and the reasoning must be reconstructable years later. A build expresses retention as rules over message attributes with effective dating and a full audit of every disposition, so that when a message is finally deleted you can show precisely which rule permitted it and confirm no hold applied. That is the difference between defensible disposal and hoping.

Problem 4: production is judged on completeness you cannot currently prove

The reason production requests are frightening is not effort, it is uncertainty. Was every channel included. Was the contractor period covered. Did the WhatsApp connector fail for eleven days in June and did anyone notice.

The feature that addresses this is unglamorous and is the one we would build first: continuous capture assurance. Every channel reports expected versus received volumes on a schedule, gaps raise an alert with an owner, and each remediation is recorded. Then a production package can carry a completeness statement covering the exact window, listing the channels in scope, the identities included and any known gaps with their explanations. Volunteering a documented, explained gap is survivable. Having a gap found by the person reviewing your production is not.

What this costs and how long it takes

A focused first release, meaning ingestion of your existing capture feeds into a unified message store, identity resolution across channels, policy defined review populations with lexicon and classifier scoring, reviewer workflow with documented sign off, and defensible search and export, runs $95,000 to $210,000 and ships in 14 to 20 weeks. A full platform adding voice transcription and review, cross channel risk scoring, legal hold and matter management, jurisdictional retention rules, capture assurance monitoring and linkage to trade surveillance cases runs $260,000 to $700,000 phased across 9 to 16 months.

What drives cost up here specifically: channel count and the fact that each capture vendor exports in its own shape; voice, which brings transcription quality, speaker separation and multilingual handling and is reliably the most expensive channel to do properly; multilingual review, since a classifier tuned on English does not transfer to Mandarin trading chat; message volume, because storing and searching years of messages with attachments at production speed is a real engineering problem; and jurisdictional spread, as each additional regime adds retention and access rules that must coexist rather than replace one another.

What holds it down: keeping your existing capture vendors and building only the supervision, identity and production layer above them. Very few firms should be writing connectors.

Build versus buy, and when the vendor is enough

Buy end to end if you are a single jurisdiction firm on email plus one chat platform, with a headcount where a principal can genuinely review a meaningful sample, and no voice obligation. Smarsh or Global Relay will cover capture, retention and review, and building would be an expensive way to arrive at the same place.

Build the layer above capture when two or more of these are true. You run more than about five channels across more than one capture vendor. Your reviewers close flagged items in bulk and everyone privately knows the review is a formality. You cannot produce a complete communication history for a named individual across all channels in under a day. You have jurisdictional retention conflicts resolved by someone's judgement rather than by a rule. Or you already run trade surveillance and cannot link a message to a trading case without manual work.

The strategic point is simple. Capture is a commodity and should be bought. Supervision is your policy, your risk appetite and your accountability, and every one of those is firm specific. A programme that outsources supervision to a vendor lexicon has outsourced the part that carries personal responsibility for a named principal, which is not a transfer any vendor contract can actually effect.

How to choose a developer for communications supervision software

Ask them how they model identity across channels over time. Effective dated identifiers bound to a person record sourced from human resources should come up unprompted. If identity is a user table, production requests will keep missing things and you will not know until someone else finds out.

Ask what they will do about immutability and disposition evidence. You want an append only store, cryptographic integrity on stored content, and a complete audit trail of every access and every deletion with the rule that authorised it. Record keeping rules allow for approaches beyond traditional write once media, but whichever route you take you have to be able to demonstrate the controls, not assert them.

Ask specifically about voice. Transcription accuracy on a noisy trading floor, speaker separation on a turret line and handling of code words and tickers are hard problems, and a team that has not done it will underestimate it by a factor you will feel.

Ask who owns the code, the classifier training data and the cloud accounts, and get it in writing before kickoff. At Digital Heroes the client owns all of it from the first commit. Your review policy expressed as code is a supervisory artefact, and if a vendor owns it you cannot fully explain your own programme when it matters most.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. The average developer spends more than 17 hours a week dealing with maintenance issues such as debugging and refactoring, and about four of those hours on 'bad code' - waste that equates to nearly $85 billion annually worldwide in opportunity cost. Source: Stripe (2018) →
  2. The median annual wage for U.S. software developers was $133,080 in May 2024, and employment is projected to grow 15% from 2024 to 2034 - a core input to any in-house build-vs-buy TCO model. Source: U.S. Bureau of Labor Statistics (2024) →
  3. In an RCT, the no-show rate was 23.5% for patients receiving a text-message reminder versus 38.1% for the control group - a 14.6 percentage-point reduction (p = 0.04). Source: Clinical Pediatrics / PubMed Central (Lin et al.) (2016) →
  4. In a McKinsey global survey of 1,259 respondents, only about 20% said their organizations excel at decision making, and just 37% said their organizations' decisions were both high quality and high in velocity. Source: McKinsey & Company (2019) →
Deepti P. · Project Manager · Lucknow

Deepti manages client software projects with a bias toward writing things down. Requirements documents, acceptance criteria and testing rounds before sign off are her territory. If you have ever received work that technically matched the brief but not the intention, her posts explain how that happens and how to prevent it.

View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.

FAQ

Frequently asked questions

How much does communications surveillance and archiving software cost to build?
A focused first release that ingests your existing capture feeds, resolves identity across channels, runs policy defined review populations with documented sign off and supports defensible export typically runs $95,000 to $210,000 and ships in 14 to 20 weeks, based on Digital Heroes delivery experience. A full platform adding voice review, cross channel risk scoring, legal hold, jurisdictional retention and capture assurance runs $260,000 to $700,000 over 9 to 16 months. Voice is consistently the most expensive channel to do properly.
Should we replace Smarsh or Global Relay or build on top of them?
Build on top. Capture connectors to WhatsApp, Bloomberg chat, Teams, Zoom and voice turrets are mature, maintained products and rewriting them would be an expensive detour. What these platforms leave thin is supervision: a generic lexicon, a queue and a sign off checkbox. The layer worth building is the one holding your review policy, your identity map, your retention matrix and your production evidence, because those are firm specific and carry personal accountability.
Does AI actually improve message review, or is it marketing?
This is one of the clearest genuine uses in compliance technology. Keyword lists cannot distinguish somebody guaranteeing to call back after lunch from a promise about performance, so reviewers learn to close in bulk and the programme becomes a ritual. A classifier scoring messages against the behaviours your policy names, running alongside the lexicon rather than replacing it, cuts reviewed volume while raising the share of reviewed items worth reading. It must stay advisory: the model prioritises, a qualified principal decides and signs.
How do we prove our capture was complete for a production request?
With continuous capture assurance rather than after the fact investigation. Each channel reports expected against received volumes on a schedule, gaps raise an alert with a named owner, and every remediation is recorded. A production package then carries a completeness statement naming the window, the channels in scope, the identities included and any known gaps with explanations. A documented and explained gap is survivable; a gap discovered by the party reviewing your production is a much worse conversation.
What is the hardest part of multi channel communications compliance?
Identity. The same person is an email address, a Bloomberg identifier, a Teams object, a mobile number, a chat account and a turret extension, and several of those change over a career or when someone returns as a contractor. If the identity map is a spreadsheet, reconstructing a full history depends on memory. Binding every channel identifier to a person record with effective dates, sourced from human resources and registration data, is what makes a request answerable in hours.
How should retention work across multiple jurisdictions?
As a matrix rather than a single policy. A message can be subject to US books and records obligations, EU or UK market abuse requirements and privacy rules simultaneously, and the practical rule is that the longest applicable retention wins while legal holds override the schedule entirely. The part most systems handle badly is evidence: when a message is finally deleted you should be able to show which rule permitted it and confirm no hold applied. That is defensible disposal rather than hope.
Do we need to capture voice calls as well as messages?
If your registered people conduct business by phone, and most do, then voice is in scope and it is the channel firms most often defer. It is also the hardest: transcription accuracy on a noisy floor, speaker separation on turret lines, tickers and code words, and multiple languages all degrade quality. Budget it as its own phase rather than as another connector, and expect review of voice to rely on transcript search plus targeted listening rather than full transcript reading.
Can communications review be linked to trade surveillance?
Yes, and the combination produces materially stronger cases than either alone, because a trading pattern accompanied by a message discussing it is a different piece of evidence. The practical design is not one merged system but a shared person identity across trading accounts and communication channels, with cases able to reference evidence from both. Build the trading side and the communications side properly first, then link, or the joined output will be noisy in both directions.
Who owns the review policy and the code if an agency builds this?
You should own the repository, the classifier training data, the review policy definitions and the cloud accounts, agreed in writing before kickoff. At Digital Heroes the client owns all of it from the first commit. Supervisory obligations attach to a named principal at your firm, not to a vendor, so a review programme you cannot open, explain and modify is one you cannot fully defend when it is examined.
What happens to my software if the agency shuts down or we stop working together?
Nothing dramatic, if the engagement was set up correctly: the code sits in your repository, hosting runs on your cloud account, and a handover document explains how to deploy and operate the system. Any competent replacement team can then take over in days rather than months. If the agency controls the repo, the servers, or the domain, fix that now, because renegotiating access during a dispute is the most expensive place to discover the problem.
Is a freelancer or an agency better for building an internal tool?
A solid freelancer works for a single-workflow tool under roughly $10,000, if you accept that one person holds all the knowledge. An agency earns its premium once the tool spans departments or integrations, because you get a developer, a designer, and a project manager plus continuity when someone leaves or gets sick. The hidden freelancer cost appears 18 months later when you need changes and the original builder has moved on, a rescue situation Digital Heroes is hired for regularly.
At what point does Retool cost more than building a custom tool?
The crossover usually lands between 25 and 50 daily users. At Retool's published Business rates of $50 per standard user and $15 per end user monthly, a 40-person deployment with a typical seat mix runs roughly $9,000 to $15,000 per year, every year, while a comparable custom tool built once for $20,000 to $30,000 carries no per-seat fees and costs about 15 to 20 percent of the build price annually to maintain. On a three-year horizon, custom comes out ahead for most growing teams in Digital Heroes engagements.
How many SaaS seats do we need before building custom becomes cheaper?
The crossover usually shows up between 20 and 50 seats on premium tiers. Salesforce Enterprise lists at $165 per user per month, so 40 users cost about $79,000 a year in subscriptions, which is real money against a custom system you would own outright. Run the comparison over three years: if subscription spend beats the build cost plus 15-20% annual maintenance, custom wins on price before you even count workflow fit.
How long does it take to build an internal tool from scratch?
A working first version typically ships in 4 to 8 weeks, and larger multi-module tools run 10 to 16 weeks. Across Digital Heroes internal tool projects the schedule splits into roughly one week of process mapping, 3 to 6 weeks of build, and 1 to 2 weeks of testing with your actual staff. The most common delay is not development but waiting on the client for sample data and workflow decisions, so name one internal owner before kickoff.
How do I calculate whether custom software will pay for itself?
Divide the build cost by the monthly benefit, where benefit is hours saved times loaded hourly cost, plus subscription fees replaced, plus any revenue the software unlocks. Three staff saving 10 hours a week each at a $40 loaded rate is about $62,000 a year, which pays back a $60,000 build in roughly 12 months. Across Digital Heroes internal-tool projects, 12 to 24 months is the normal payback range, and anything projecting under 6 months usually means the spreadsheet is hiding costs.
Can a custom internal tool connect to QuickBooks, Salesforce, and the other software we already use?
Yes, and integrations are usually the strongest argument for going custom instead of chaining tools together with Zapier. QuickBooks, Salesforce, Shopify, Stripe, Slack, and Google Workspace all have mature APIs, and each integration typically adds $1,500 to $5,000 to a Digital Heroes build depending on how much two-way syncing you need. The honest caveat is legacy industry software without an API, which may need file-based imports instead of a live connection, so list every system in the first conversation.
How small can the first version of my software be and still be worth building?
One workflow, end to end, for one type of user: the single process that currently burns the most hours or loses the most money. In Digital Heroes delivery experience, first versions scoped to 6 to 10 weeks of build time ship, get used, and generate the feedback that makes version two obviously right, while 9-month first versions routinely launch with features nobody touches. Everything you cut from v1 gets cheaper to build later, because real usage reorders the roadmap for you.
What tech stack should an internal tool be built with?
Boring and popular: a React or Next.js frontend, a Node.js or Python backend, and PostgreSQL covers the vast majority of internal tools and keeps future hiring easy. The stack matters far less than whether a different developer can pick the code up in two years, so require documentation as a deliverable and avoid anything exotic. Treat it as a red flag if an agency pushes a proprietary platform only they maintain, because that quietly converts your tool into a subscription to that agency.
Should we build our internal tool in Retool instead of hiring developers?
Retool is the right choice if someone on your team is comfortable with SQL and JavaScript and the audience is a handful of technical users, because a basic CRUD dashboard comes together in days. Hire developers when non-technical staff will use the tool daily, when the logic goes beyond forms sitting on a database, or when per-seat pricing stings, since Retool's Business tier lists at $50 per standard user per month. A pattern Digital Heroes sees often: companies arrive after a year on Retool with a tool nobody can maintain because the one person who built it has left.
Who owns the code when an agency builds our internal tool?
You should, outright, with full IP transfer in the contract and the code delivered to a repository you control, such as your own GitHub organization. Digital Heroes transfers complete ownership on final payment as standard practice, and any agency that keeps the code or licenses it back to you is building a dependency you will pay for later. Confirm you also own the hosting, domain, and database accounts, since many of the vendor disputes Digital Heroes gets called into involve infrastructure registered under the agency's name.
Who can build a custom internal tools system?

Digital Heroes builds custom internal tools systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other internal tools companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading
let's build

Build something worth launching.

A plan, a team, a timeline, within 24 hours. No decks, no discovery calls. Tell us what you're building and we'll come back with a real scope and a real number.

message us directly · we reply within one business day

mission briefing

Monthly dispatch

Playbooks, real build costs, and what we're shipping. One email a month. No fluff.

visit us

New York HQ

1140 Broadway, Suite 704 · New York, NY 10001

Get directions
Online now

Hey there 👋 How can we help you today?