Industry guide · Custom Software

AML Transaction Monitoring: Building a System That Can Defend Its Thresholds to an Examiner

Aml Transaction Monitoring software visual showing radar, funnel, and sliders horizontal.
The short answer

If your alert volume is unmanageable and you cannot show an examiner why a threshold is set where it is, the fix is usually not a new engine. A first release covering a monitoring data layer, scenario execution with full parameter versioning, and alert triage with disposition capture runs $100,000 to $240,000 and ships in 14 to 20 weeks in Digital Heroes delivery experience. A full platform adding segmentation, above and below the line testing, tuning evidence packs and model documentation runs $280,000 to $700,000 over 9 to 18 months. If you are a community bank with conventional retail and commercial products, buy Verafin or a comparable engine and build only the data quality and tuning evidence layer around it. Replacing a scenario library you did not need to replace is the most expensive mistake in this category.

Why monitoring programs fail the exam and not the criminals

An examiner rarely asks whether your system caught a launderer. They ask why your structuring scenario fires at a particular aggregate over a particular number of days, who approved that, what analysis supported it, what the alert and case outcomes were at that setting, and what happened when you tested just below the line. That is four questions about evidence and one about detection.

In most institutions the answers live in three places. The threshold is in a vendor console. The justification is in a slide deck from a consultant engagement two years ago. The outcomes are in a case system that does not share keys with the alert system. Nobody can assemble those into one document without a project, which is why the finding is so often written about governance rather than detection.

The second failure is quieter. Alerts pile up, the queue grows, and the response is to raise thresholds because the team cannot cope. That is a capacity decision being made as a risk decision, and it is undocumented by definition. Everyone in financial crime knows this happens. Almost no monitoring system is built to make it visible.

Problem one: the data layer is the actual weakness

Scenario logic is comparatively simple. Aggregate cash activity by customer across a rolling window. Compare wire activity to expected activity captured at onboarding. Flag rapid movement through an account. The reason these produce noise is almost never the logic. It is that the customer is three customer records because the core, the card processor and the digital channel each created one. It is that expected activity was collected once at account opening in a free text field. It is that a counterparty name arrives differently from each channel, so two hundred transfers to the same beneficiary look like two hundred beneficiaries.

NICE Actimize and Oracle Financial Crime and Compliance Management ship deep scenario libraries and are the conservative, examiner familiar choice, and that matters. The honest limitation is that tuning is slow and costly through a vendor cycle, and explaining behaviour you cannot inspect is harder in a room with an examiner. Verafin is genuinely strong for community banks and credit unions and brings cross institution context most single institutions cannot build, provided your profile matches the one it was designed around. Feedzai and Hawk bring modern detection with better false positive characteristics, and the burden then shifts to explainability and validation under model risk expectations, which is real work rather than a reason to avoid them.

None of them can fix your entity resolution. That is your data, your channels and your history, and it is where the first fifty thousand dollars of a monitoring project should go, whatever engine you keep.

Problem two: thresholds without lineage are indefensible

What a defensible build looks like: every scenario has a version, every parameter change has an author, a date, a rationale and an approval, and the system can reconstruct exactly which scenario version and parameter set produced any historical alert. When an examiner asks why the threshold moved in March, you produce the change record, the analysis that supported it, the approver, and the alert and case outcomes for the ninety days either side.

Add above the line and below the line testing as a native function rather than a consulting exercise. Below the line sampling means taking activity that fell just under the threshold, sampling it, and having investigators review it to demonstrate you are not missing productive alerts. Above the line means testing whether raising a threshold would have lost real cases. Both are standard practice, both are usually done manually once a year by an outside firm, and both become routine when the system can replay historical transactions through a candidate parameter set. That replay capability is the single feature that changes the character of a monitoring program, because it turns tuning from an argument into a measurement.

Problem three: your products are not in the scenario library

Vendor scenarios encode conventional banking. If you are a conventional bank, that is a strength. If you run a money services business portfolio, a payments company with sub merchants, a digital asset on ramp, trade finance, a banking as a service programme with fintech partners, or a cannabis related banking program, the typologies that matter to you are not in the box, and the vendor customisation to add them costs more than building them.

The FFIEC examination manual expects your monitoring to reflect your risk assessment. If your risk assessment names a typology your monitoring cannot express, that gap is written down in your own documents, which is the worst possible place for it to be. This is the strongest genuine case for custom scenario work: not replacing an engine, but adding the detection your specific products require, with the same parameter versioning and testing discipline as everything else.

Problem four: alerts, cases and customer risk do not share a spine

An investigator opens an alert, then opens the core to see the customer, then opens the KYC system to see the risk rating and the expected activity, then searches the case system for prior alerts on the same customer, then checks a negative news tool. Five systems, no shared identifier, twenty minutes before any thinking starts.

What a custom build does: one customer spine that every alert, case, transaction and risk rating hangs from. The alert opens with the customer's profile, the expected activity captured at onboarding and at last review, prior alerts and their dispositions, related parties, and the specific transactions that triggered the scenario with the aggregate shown. Disposition capture is structured, not free text, because the pattern of dispositions is the raw material of tuning. Two hundred alerts closed with the reason recorded as expected business activity for a customer segment is a tuning insight. The same two hundred closed with a free text note is nothing.

Where AI genuinely helps and where it does not

Two places earn their keep. Entity resolution across channels and counterparty name normalisation, which is a matching problem that machine learning is well suited to and that no rule set solves. And alert triage scoring, used to order the queue rather than to close alerts, so investigators spend their first hours on the alerts most likely to be productive. Scoring the queue is defensible because nothing is suppressed.

Where we advise caution: using an opaque model to close alerts automatically. Under model risk expectations you will be asked to validate and explain it, and a model that cannot be explained becomes an examination finding rather than an efficiency. If you go there, budget for independent validation and full documentation from the start, and keep a rules based safety net for the typologies your risk assessment names explicitly.

What it costs and how long it takes

A first release, meaning the customer and transaction data layer with entity resolution, scenario execution with parameter versioning, and alert triage with structured dispositions, runs $100,000 to $240,000 and ships in 14 to 20 weeks. A full platform adding customer segmentation, historical replay for above and below the line testing, tuning evidence packs, model documentation and integration with case management and filing runs $280,000 to $700,000 across 9 to 18 months.

What pushes cost up specifically: the number of source systems feeding transactions, since each channel is a separate mapping and a separate data quality argument. Historical depth, because replay testing needs several years of transactions in a queryable shape. Product complexity, particularly correspondent banking, trade finance and partner banking programmes where the customer of your customer matters. Real time requirements, if some scenarios must decision inside a payment path rather than overnight. And the state of your KYC data, which is usually worse than the compliance team believes and always takes longer to remediate than the plan says.

Build versus wrap versus buy

Buy the engine if you are a community bank or credit union with conventional retail and commercial products. Verafin and its peers cover your typologies and your examiner has seen them before. Your money goes into data quality and tuning evidence, not into rewriting scenarios someone else has already validated.

Wrap the engine, which is the answer we give most often, when the engine is adequate but the surrounding evidence is not. Keep the vendor detection, build the data layer, the parameter governance, the replay testing and the tuning documentation around it. This is typically a third of the cost of a replacement and addresses what the finding actually said.

Build detection yourself when your products are genuinely outside the library: payments companies with sub merchant flows, digital asset businesses, banking as a service programmes, money services businesses with corridor specific typologies. In those cases the vendor scenarios are approximations of somebody else's business, and you will pay for customisation that is worse than the thing you could have owned.

How to choose a developer for financial crime work

Ask how they will reconstruct which scenario version and parameter set produced an alert from eighteen months ago. If the design does not version parameters and store the executing version against the alert, the system cannot be defended and should not be built.

Ask how they intend to handle entity resolution across your channels, and listen for whether they ask about your core, card processor and digital channel identifiers in the same breath. A team that treats the customer as a given has not worked on monitoring.

Ask what they will do about below the line testing. If it is not a function of the product, you will be paying a consultant to do it manually every year forever.

Ask who owns the code, the scenario definitions and the tuning history, and put it in the contract before kickoff. Your tuning history is your regulatory defence and it compounds. At Digital Heroes the client owns all of it from the first commit, and we would tell you to walk away from anyone whose contract makes your threshold justifications their property.

Research & sources

The evidence behind this guide

Independent findings on why this investment pays off. Every link goes to the primary source.

  1. Deloitte's research found that digitally advanced small businesses experienced revenue growth nearly 4x as high as the prior year, were about 3x as likely to have exported, were nearly 3x as likely to have created new jobs, and were more than 3x as likely to have seen more sales inquiries in the last year. Source: Deloitte (research summarized by Google) (2017) →
  2. Across more than 5,400 IT projects studied by McKinsey and the University of Oxford BT Centre, large IT projects ran on average 45% over budget and 7% over schedule while delivering 56% less value than predicted. Source: McKinsey & Company / University of Oxford (BT Centre for Major Programme Management) (2012) →
  3. In the Flexera 2025 State of ITAM report, respondents reported roughly 33% of SaaS spend is wasted, underscoring how paying for off-the-shelf seats and tiers that go unused erodes the supposed cost advantage of generic SaaS. Source: Flexera (2025) →
  4. Retailers connecting point-of-sale and loyalty data in an omnichannel strategy reported up to 15% lower cost per purchase and nearly 20% higher incremental store revenue. Source: Deloitte (2024) →
Shreyansh S. · Managing Director · Lucknow

Shreyansh runs the Lucknow operation, sitting between clients who need software built and the teams who build it. Most of his week goes on scoping work honestly, deciding what a project should and should not include, and keeping delivery promises realistic. He writes for readers weighing up whether to commission custom software at all.

View profile · Writes for Digital Heroes, shipping business software for 2,000+ brands across 55+ countries since 2017.

FAQ

Frequently asked questions

How much does it cost to build AML transaction monitoring software?
A first release with a monitoring data layer including entity resolution, scenario execution with parameter versioning and alert triage with structured dispositions runs $100,000 to $240,000 and ships in 14 to 20 weeks in Digital Heroes delivery experience. A full platform adding segmentation, historical replay testing, tuning evidence and model documentation runs $280,000 to $700,000 over 9 to 18 months. Source system count and the state of your KYC data drive the range more than asset size does.
Should we replace NICE Actimize or build around it?
Wrapping is the answer we give most often. If the finding was about governance, tuning evidence or data quality rather than missed typologies, keep the vendor detection and build the data layer, parameter governance, replay testing and documentation around it, typically for a third of the cost of replacement. Replace or build detection only when your products genuinely sit outside the scenario library, such as sub merchant payment flows, digital assets or banking as a service programmes.
Why do we get so many AML false positives?
Almost always because of data rather than logic. The same customer exists as three records across the core, the card processor and the digital channel, expected activity was captured once at account opening in a free text field, and counterparty names arrive differently per channel so one beneficiary looks like two hundred. Fixing entity resolution and expected activity capture reduces alert volume more than any threshold change, and unlike a threshold change it does not reduce coverage.
How do we make threshold tuning defensible to examiners?
Version every scenario and every parameter change with an author, date, rationale and approver, and store the executing version against each alert so any historical alert can be reproduced. Then make above the line and below the line testing a native function by replaying historical transactions through candidate parameter sets, rather than commissioning it once a year. When the question comes, you produce the change record, the supporting analysis and the outcomes on either side of the change.
Can machine learning replace rules in transaction monitoring?
It can improve two things safely: entity resolution across channels, and scoring the alert queue so investigators reach productive alerts first. Using an opaque model to close alerts automatically is where institutions get into trouble, because you will be asked to validate and explain it under model risk expectations, and a model you cannot explain becomes a finding rather than an efficiency. If you go that route, budget for independent validation and keep a rules based safety net for the typologies your risk assessment names.
What does below the line testing actually involve?
You sample activity that fell just under a scenario threshold, have investigators review it as if it had alerted, and measure whether productive cases were being missed at the current setting. Above the line testing is the mirror image: whether raising a threshold would have lost real cases. Both are standard practice and both are usually done manually by an outside firm once a year, which is why building historical replay into the system changes the economics of tuning entirely.
How long does an AML monitoring build take?
Fourteen to twenty weeks for a first release, and nine to eighteen months for a full platform with replay testing and documentation. The long pole is rarely scenario development. It is mapping each channel's transaction data into one model, resolving customers across systems, and remediating KYC data that is usually in worse shape than the compliance team expects. Institutions with a single core and clean customer identifiers move considerably faster.
Our products are not in any vendor scenario library. What then?
That is the strongest genuine case for custom detection. Money services corridors, sub merchant payment flows, digital asset on ramps, trade finance and partner banking programmes carry typologies that vendor libraries approximate rather than express. Since examination expectations tie monitoring back to your own risk assessment, a typology named in your risk assessment that your system cannot detect is a gap documented in your own files. Build those scenarios with the same versioning and testing discipline as the rest.
Who owns the tuning history if a vendor or agency builds our system?
You should own the repository, the scenario definitions, the parameter history and the cloud accounts, written into the contract before kickoff. Your tuning history is your regulatory defence and it becomes more valuable every year as outcomes accumulate against it. At Digital Heroes the client owns all of it from the first commit, and any arrangement that makes your threshold justifications someone else's property should be refused.
How do I make sure custom software is secure and compliant with rules like HIPAA?
Start with the baseline every business system should have: encryption in transit and at rest, role-based access control, and audit logs. If HIPAA applies, the hosting provider must sign a Business Associate Agreement, which AWS, Azure, and Google Cloud all offer, and access controls have to be designed in from day one, not bolted on. SOC 2 certifies a company's operating practices, not a codebase, so ask vendors what they have shipped in your regulated domain rather than which logos are on their website.
Can I build my product on a no-code tool like Bubble instead of hiring developers?
For testing whether anyone wants the product, yes, and Bubble's paid plans start at $29 a month, which is the cheapest validation you will ever buy. The ceiling arrives with complex data relationships, heavy integrations, performance at a few thousand users, and the fact that you cannot export a Bubble app to servers you control. A path many Digital Heroes clients take: prove demand on no-code, then rebuild custom once revenue justifies it, treating the no-code version as a paid prototype rather than a foundation.
Can we migrate years of data out of our current system into new custom software?
Almost always yes, through CSV exports or the vendor's API, and migration should be scoped as its own workstream with field mapping, a dry run, and a planned cutover window rather than an afterthought. The real time sink is rarely moving the data; it is cleaning it, since years of duplicates, free-text fields, and inconsistent formats surface all at once. Pull a full export from your current vendor before committing to anything new, because some SaaS plans restrict exports on lower tiers.
Will an app built for 10 users survive growing to 500?
Yes, if it is built on standard cloud infrastructure with a sound data model, because moving from 10 to 500 users is a hosting configuration change, not a rebuild. The scaling decisions that actually hurt are made early and invisibly: how the database is structured, how accounts and permissions are modeled, and whether background work is queued properly. Ask your agency how the system would handle ten times the load; the right answer is boring and specific, and a promise to cross that bridge later means you will pay for the bridge twice.
How many SaaS seats do we need before building custom becomes cheaper?
The crossover usually shows up between 20 and 50 seats on premium tiers. Salesforce Enterprise lists at $165 per user per month, so 40 users cost about $79,000 a year in subscriptions, which is real money against a custom system you would own outright. Run the comparison over three years: if subscription spend beats the build cost plus 15-20% annual maintenance, custom wins on price before you even count workflow fit.
What is a discovery phase, and is it worth paying for separately?
Pay for it, and treat the output as yours. A discovery phase runs two to three weeks, typically 5 to 10% of the eventual build budget, and produces a written scope, wireframes, and a fixed quote you can take to any vendor, including a competitor of the agency that wrote it. Skipping it is how projects end up quoted from a two-paragraph email and delivered at twice the price.
Couldn't I just build my app in Bubble or another no-code tool instead of hiring an agency?
For validating an idea with real users, yes, and we tell clients that honestly. The walls come later: Bubble apps cannot be exported as code to run anywhere else, performance drops on complex data operations, and usage-based pricing climbs as you grow. A meaningful share of Digital Heroes custom builds are rebuilds of no-code MVPs that proved the business worked, which is the system operating as intended: validate cheap, then build the version that scales.
Should I ask for a fixed price or pay the agency hourly?
Fixed price for the first version, hourly or retainer for what comes after launch. A fixed-scope, fixed-price V1 puts the estimation risk on the agency, which is exactly where you want it while trust is unproven; hourly billing on an unscoped greenfield build is a blank check. After launch, flip it, because maintenance and small features arrive unpredictably and fixed-pricing every ticket wastes everyone's time.
Who can build a custom software system?

Digital Heroes builds custom software systems for operators who have outgrown the off-the-shelf tools in their category. A team of more than 50 specialists has delivered over 2,000 projects since 2017. Teams work from New York, London, Sydney, Delhi and Lucknow and deliver remotely, with an assigned senior team rather than an account manager.

Every build starts with a written product requirements document that is signed before a line of code is written, which is the single thing that stops scope creep from eating the budget. Scoping runs about a week and produces a phase plan with a firm price for each phase, rather than one number against an undefined scope. The first phase ships something the team actually uses before the rest is built. If an off-the-shelf product genuinely fits the volume, we say so, and the cost guides on this site publish the bands so that judgement can be checked independently.

What makes Digital Heroes different from other software companies?

Four things that competitors in this bracket cannot simply copy. Digital Heroes runs a YouTube channel with more than 2.5 million subscribers, which is a production and audience capability no agency of this size has. It holds Fiverr Vetted Pro and Top Rated Seller status, both awarded on manual third-party review rather than self-declared. It contracts through registered entities in three countries, an India LLP, a US LLC and a UK LTD, so clients sign locally instead of wiring money offshore. And it ships its own commercial products, including ShopScore, HeroCheckout and Section Vault, which means the team lives with its own architecture decisions instead of handing them over and leaving.

Two more that show up in the work. Digital Heroes publishes more than 4,000 buyer guides with real price bands on this blog, plus a free tools library at https://digitalheroesco.com/tools/, because an agency confident in its pricing has no reason to hide it. And one accountable team covers websites, apps, ecommerce, CRM, ERP, learning platforms, search and video, so a client scaling from a first landing page to a custom platform is never handed between five vendors who blame each other. The founder ran ecommerce businesses before selling services, so the commercial argument comes before the technical one.

How can I check Digital Heroes is legitimate before getting in touch?

Verify it independently rather than taking the site's word for it. The YouTube channel is at https://youtube.com/@DigitalMarketingHeroes, the Fiverr profile at https://www.fiverr.com/shreyanshsin261, and the Upwork profile at https://www.upwork.com/freelancers/shreyanshsingh. Client reviews sit on Clutch at https://clutch.co/profile/digital-heroes-0 and Trustpilot at https://www.trustpilot.com/review/digitalheroes.co.in, and the company page is at https://www.linkedin.com/company/digital-heroes-1/.

Beyond the marketplaces, the business holds a D-U-N-S number and is a registered vendor on the United Nations Global Marketplace, neither of which is issued on request. Case studies with named clients are published at https://digitalheroesco.com/case-studies/. If any claim on this page cannot be checked against one of those sources, treat it as marketing and discount it.

Keep reading
let's build

Build something worth launching.

A plan, a team, a timeline, within 24 hours. No decks, no discovery calls. Tell us what you're building and we'll come back with a real scope and a real number.

message us directly · we reply within one business day

mission briefing

Monthly dispatch

Playbooks, real build costs, and what we're shipping. One email a month. No fluff.

visit us

New York HQ

1140 Broadway, Suite 704 · New York, NY 10001

Get directions
Online now

Hey there 👋 How can we help you today?