eDiscovery Software Problems: The 7 That Cost Real Money, and How to Avoid Them
The most expensive failure is scoping a review platform when what the firm needed was everything around one. Teams set out to replace Relativity, spend a year building coding panels and search, and arrive with something slower and less capable than the product they already licence, while the actual leak stays open: legal hold still runs on a spreadsheet, collections still have no chain of custody, and the promotion decision that sets your hosting bill still happens by default under time pressure. The money goes on rebuilding a commodity and the per gigabyte invoice that provoked the project never moves.
Why does the rebuild Relativity trap catch so many firms?
Because the pain is felt in review, so review is what gets scoped. Reviewers complain about the interface, partners complain about the hosting invoice, and both complaints point at the same screen. The natural conclusion is that the platform is the problem.
It usually is not. Relativity, Everlaw and DISCO represent years of hardening, analytics and scale, and Nuix has a processing engine that handles difficult formats better than most. Reproducing any of that inside a sane budget is not realistic, and a project that tries will spend its entire allowance on coding panels, search and viewer performance, which are exactly the parts already solved.
Meanwhile the decisions that determine cost happen before anything reaches review. What gets collected, what gets promoted into hosted storage and how long it stays there are set in the first fortnight of a matter, by people who cannot see the financial consequence at the moment they choose. Nobody makes a bad call. The default is simply to host everything, because hosting everything is the fastest route to reviewing anything.
The fix is to draw the boundary in the statement of work before design starts. Build legal hold and custodian tracking, collection orchestration with chain of custody, pre hosting assessment that culls, and a live matter cost model. Keep your review platform. The one genuine exception is a residency requirement where client data cannot leave your infrastructure, and that should be stated as a constraint on page one rather than discovered as a preference later, because it changes the architecture and the number substantially.
What goes wrong when chat and linked documents are collected?
The hard part of a modern collection is not the mail archive. It is Slack and Teams, and the failure is not technical so much as definitional.
A chat channel has no natural document boundary. Somebody has to decide what a reviewable unit is, and if that decision is made implicitly by an export script rather than explicitly by a lawyer, you have adopted a position you cannot defend at a meet and confer. The common defensible answer is a channel and day boundary with participants preserved, but the point is that it is written down, applied consistently and disclosable.
Linked documents are the sharper trap. A message references a document by link rather than attaching it, so the family relationship every reviewer relies on does not exist in the data. The version you collect is the version at collection time, which is not necessarily what the recipient saw. Builds routinely resolve links where the tenancy permits and simply skip the rest, producing a corpus with a silent gap.
Three rules prevent the expensive version of this:
- The review unit rule is documented, versioned and exportable, so you can state exactly how chat was processed and when the rule changed.
- Linked documents are resolved where the tenancy allows and logged where they cannot be, with the reason. A documented gap is defensible. A silent one is not.
- The collection records what was attempted and failed, not only what succeeded, because absence of evidence in a collection log reads as absence of diligence.
Anyone who proposes exporting a channel to a document format and calling it collected has not done this work and will hand you a production problem two months later.
Why do the source system integrations break after launch?
Because each source is a separate product with its own release cycle, and collection interfaces change more often than anyone budgets for.
Microsoft 365, Google Workspace, Slack, Teams and mobile forensic tooling are five integrations with five different export shapes, five different permission models and five different throttling behaviours. Throttling is the one that catches teams out. A collection that ran cleanly in testing against a small tenancy will hit rate limits on a real one, and if the orchestration treats a throttled response as a failure it will retry until something gives up, quietly leaving part of a custodian uncollected.
Permission changes are the other recurring break. Collection runs under a service identity, and a tenancy administrator tightening scopes during an unrelated security review can remove access to exactly the mailboxes you need. The symptom is a collection that completes and returns less.
Design for both. Every collection produces a manifest stating what was requested, what was returned, what was throttled and what failed, with hashes for what was acquired, and the manifest is compared against the request rather than assumed to match. A collection is not complete because the job finished. It is complete because the counts reconcile and a named person accepted the exceptions. That manifest is also your chain of custody evidence, so building it well solves two problems at once.
What happens when legal hold and production gates are underbuilt?
These are the two places where an underbuild turns into a sanctions conversation rather than an inconvenience.
A hold is not a notice, it is an ongoing obligation, and spreadsheet tracked holds fail in predictable ways: a custodian who never acknowledged and was never chased, a retention policy that kept running because nobody told the systems team, a departing employee whose device was wiped by standard offboarding. Under Rule 37(e) the consequences of losing electronically stored information that should have been preserved can be serious, and the first question is always what steps you took. If the build models a hold as a sent email with a status field, it produces no answer to that question.
Production is the mirror image. The output has a form set by agreement or order: image format with load files, natives where required, a numbering scheme that must not collide across volumes, redactions burned correctly, confidentiality designations applied under the protective order, and a privilege log that describes withheld documents adequately under Rule 26(b)(5) without waiving what it protects. Mistakes here are public: a redaction applied as a rectangle over selectable text, a privileged parent produced because an attachment was tagged and the parent was not, a volume overlapping a prior number range.
The fix in both cases is to replace diligence with gates. Holds are objects with scope, custodians drawn from your directory so leavers are detected automatically, evidenced reminders and a recorded release at matter close. Productions cannot be released until automatic checks pass for text under redactions, family completeness, privilege consistency across every member of a family, numbering continuity and load file integrity. And every production is reproducible, so when the other side queries volume three you rebuild exactly what was sent rather than reconstructing it.
Should you build custom or configure what you already own?
Buy the review platform. For a firm handling ordinary matter volumes, Relativity with disciplined process beats any build, and Everlaw or DISCO are the right answers where a simpler commercial model and usability matter more than extensibility. Nuix is the answer if difficult processing is your actual bottleneck rather than review.
More importantly, before commissioning anything, find out what your existing platform already does. Relativity in particular is extensible, and a large number of firms run it with default workflows and no scripted validation because the implementation happened years ago and nobody revisited it. Threading, deduplication across custodians and search term reporting are frequently available and frequently unused. If your culling is poor because nobody switched it on, a custom build will not fix a habit.
Build the surrounding layer when two or more of these hold. Your hold tracking is a spreadsheet and a preservation question has already been raised against you. You cannot tell a client the cost consequence of a scoping decision at the moment the decision is made. You handle matters where data cannot leave a jurisdiction. Your collections span several systems and your chain of custody is assembled from emails afterwards. Or you are a corporate department managing several outside firms and need one cost and protocol model across all of them, which no single firm's platform will ever give you.
How do hidden costs get into the quote?
Four items account for most of the overruns, and none of them is a surprise if you ask early.
Source systems counted as one line. Five collection sources are five integrations with five export shapes and five throttling behaviours. Price them individually and name them in the contract.
Scale treated as configuration. Text extraction and search across terabytes is genuine engineering, not a deployment setting. A quote that does not state an assumed volume has not been priced against yours.
Residency discovered late. If client or regulatory requirements mean data cannot leave a jurisdiction, that constraint drives architecture and cost more than any feature. Establish it before design, not during.
Security review and deletion at matter close. Holding client litigation data after the obligation ends is a liability rather than a service, so the deletion process, key handling and a security review belong in the build budget. Retrofitting them is the sort of thing that surfaces during a client audit at the worst moment.
What separates a build that works from one that fails here?
Ask them what a family is and what happens when a parent email is privileged and an attachment is not. If that needs explaining, they will build a document management system with a legal skin and you will find out at your first production.
Ask how they would handle a channel with forty thousand messages and links to cloud documents. The answer should include a documented review unit rule, resolution of linked documents where the tenancy permits, and logging where it does not. Anyone who says they will export it to a document format has not done this.
Ask what pre production validation they will build, and expect text under redaction, family completeness, privilege consistency across families, numbering continuity and load file integrity to be automatic gates rather than a checklist someone runs. That is the difference between a defensible process and a diligent person who was on holiday.
Ask how the cost model gets in front of a partner at the moment of the promotion decision, because a partner looking at a screen showing what a custodian set adds per month scopes differently from one told the total six weeks later. Then settle code ownership, hosting location and encryption in writing before kickoff, with a documented deletion process at matter close. At Digital Heroes the client owns the repository and the infrastructure accounts from the first commit.
The evidence behind this guide
Independent findings on why this investment pays off. Every link goes to the primary source.
- Technology 'Leaders' grow revenue at more than twice the rate of 'Laggards'; laggards surrendered 15% in foregone annual revenue in 2018 and stood to miss out on as much as 46% in revenue gains by 2023 if they did not change their enterprise technology approach. Based on a survey of more than 8,300 organizations across 20 industries and 20 countries. Source: Accenture (2019) →
- The performance gap between digital and AI leaders and laggards is widening: McKinsey reports leaders pull ahead on shareholder returns, and the average maturity spread between top and bottom performers jumped ~60% (from 10 points in 2016-19 to 16 points in 2020-22), reinforcing that the returns to transformation concentrate among top performers. Source: McKinsey & Company (2023) →
- APQC's Open Standards Benchmarking data on the monthly financial close found median performers take about 6.4 calendar days to close the books, while top performers (top 25%) do it in 4.8 days or fewer and bottom performers (bottom 25%) take 10 or more days. Source: APQC (2018) →
- SMS reminders that stated the specific cost of the appointment to the health system reduced missed appointments in Trial One, with the DNA (did-not-attend) rate falling from 11.1% (control) to 8.4% (specific-costs message) - an odds ratio of 0.74 (95% CI 0.61-0.89), i.e. roughly a 24-26% relative reduction - at no additional cost. (Trial Two replicated this at an 8.2% DNA rate.). Source: PLOS ONE (Hallsworth et al.) (2015) →
Oliver runs UK client accounts day to day, chairing the calls where scope, budget and timeline meet reality. He is useful reading for anyone about to commission custom software and wondering what a healthy agency relationship should feel like from the client side.
View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.
Frequently asked questions
Should we build our own review platform instead of using Relativity?
How do we cut per gigabyte hosting without cutting corners?
How should Slack and Teams be turned into reviewable documents?
Why do collections come back incomplete without any error?
What makes a legal hold defensible in software?
Which production checks should be automatic gates?
What gets underpriced in eDiscovery build quotes?
How do we get partners to scope matters differently?
What should I prepare before contacting a software development agency?
Is a solo freelancer enough for my project, or do I really need an agency?
What happens to my software if the agency shuts down or we stop working together?
What happens if I stop paying for maintenance after launch?
Who owns the code when an agency builds my software?
How do we get years of data out of our old system and into the new one?
How many SaaS seats do we need before building custom becomes cheaper?
Will custom software work with the tools we already use, like QuickBooks and Stripe?
Does it matter which tech stack the agency wants to use?
Should I hire a freelancer or an agency for my software project?
Is it cheaper to customize Salesforce than to build a custom CRM from scratch?
Can I build my product on a no-code tool like Bubble instead of hiring developers?
Who can build a custom software system?
Digital Heroes builds custom software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.
Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.
What makes Digital Heroes different from other software companies?
Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.
Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.
How can I check Digital Heroes is legitimate before getting in touch?
Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.
Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.