Transaction monitoring is the engine of an AML programme's detection capability — but a poorly calibrated engine generates excessive alerts that strain analyst resources, and that's not just an efficiency problem. When a team spends most of its time clearing obvious non-issues, genuine suspicious activity risks getting missed or under-investigated in the noise. AUSTRAC, the FCA, and FinCEN have all identified poor transaction monitoring calibration as a key driver of adverse examination findings.
Why do false positive rates tend to rise over time?
Because monitoring rules typically start from generic industry thresholds rather than institution-specific data, and then stay static while the customer base and product mix keep changing underneath them. A rule configured for 50,000 retail customers five years ago can be severely miscalibrated for a current portfolio of 500,000 customers that now includes substantial business banking. Regulatory pressure to demonstrate coverage after an examination often compounds this — institutions add new monitoring rules without retiring old ones, creating overlapping rules where a single transaction triggers multiple alerts without adding any real detection value.
What does a structured approach to tuning actually look like?
Start by analysing each rule's alert-to-report conversion rate. A rule converting fewer than 1 in 100 alerts into an actual Suspicious Matter Report is a strong candidate for threshold adjustment, retirement, or a logic redesign — that low a conversion rate signals the rule isn't finding real risk, just generating noise. Threshold calibration should be data-driven: using historical transaction data to model alert volumes at different settings and understand the sensitivity-specificity tradeoff at each one, with the analysis documented, since regulators expect threshold choices backed by data rather than arbitrary adjustment. Peer comparison offers useful context, but it can't substitute for institution-specific analysis — a genuinely risk-based approach requires parameters that reflect the institution's own actual profile, not an industry average.
How does customer segmentation help reduce false positives?
By replacing a single uniform threshold with thresholds calibrated to each customer type's actual expected behaviour. Applying the same threshold across retail individuals, small businesses, corporates, and private banking generates excessive false positives in lower-risk segments while potentially missing genuinely elevated patterns in higher-risk ones — a $50,000 transaction should flag very differently for a retail account than for a corporate treasury account moving that amount routinely. Segment-specific rules materially improve the signal-to-noise ratio, reducing false positives without weakening detection sensitivity where it actually matters.
Why does documentation matter so much here?
Because regulators don't accept undocumented threshold changes — even when the change genuinely reduces false positives and improves programme efficiency. Documentation should cover: the rule being adjusted, the current and proposed thresholds, the supporting data analysis, the expected impact on alert volume and conversion rate, compliance officer sign-off, and the implementation date. Treating monitoring rules with the same governance rigour applied to credit risk models is what actually meets the standard regulators are examining against — tuning isn't a one-time activity, it's a continuous discipline that requires data analysis, governance, and documentation working together, not any one of the three on its own.



