Reducing AML false positives without missing real risk means testing a threshold on both sides, not just watching the alert count fall. Above-the-line testing checks whether a rule's alerts are worth an analyst's time. Below-the-line testing checks whether activity just under the threshold should have alerted. A defensible programme runs both and keeps a record an examiner can follow.
Most teams tune by instinct: raise a threshold until the queue feels manageable, and move on. That produces a quieter queue and a program that cannot show its work. This piece covers the testing method itself, how to document it the way BSP's own guidance expects, and where AI-assisted tuning fits and where it doesn't yet.
What is AML threshold tuning, and how is it different from calibration?
Threshold tuning is the act of adjusting a rule's parameters, a peso amount, a transaction count, a time window, so it separates genuine risk from ordinary activity more accurately. Calibration is the broader discipline around it: testing a threshold before and after a change, documenting why it moved, and reviewing it on a schedule rather than only when the queue gets loud. Tuning is the adjustment. Calibration is the evidence that the adjustment was justified.
The distinction matters because a compliance head asked “have you tuned your rules?” can usually say yes. Asked “can you show me the calibration record for this rule?” is a different question, and it's the one an examiner tends to ask.
Why does AML false positive reduction depend on testing both sides of the threshold?
Because a threshold only ever has two failure modes, and tightening it to fix one worsens the other. Set it too loose, and the queue fills with alerts that lead nowhere, the classic false positive problem, and the reason “AML false positive reduction” is one of the most searched problems in this category. Set it too tight in response, and the queue gets quieter while genuine risk passes through unflagged, a false negative that no one sees because there's no alert to review.
Above-the-line testing addresses the first failure mode directly: it samples the alerts a rule already generated and measures what share led to an escalation, an STR, or another productive outcome. A rule with a 3% productive-alert rate is not calibrated. It's noise with a threshold attached.
Below-the-line testing addresses the second, and it's the step most programmes skip, because it requires deliberately looking at activity the system chose not to flag. It samples transactions that fell just under the current threshold and checks, manually or against a known-risk sample, whether any of them resemble the typology the rule is meant to catch. If they do, the threshold is too tight, and the queue's quiet is bought with blind spots.
What does a BSP-regulated institution need to be able to show an examiner?
More than a threshold value and a change date. BSP's Memorandum M-2023-013, the Guidance Paper for an Effective AML/CTPF Transaction Monitoring System, calls out independent testing of the monitoring system as a distinct expectation, separate from having rules that run. That means a calibration record, not a verbal assurance that tuning happens.
This builds on the AML/CFT framework the BSP has developed since Circular No. 706, Series of 2011, which first set out detailed AML rules and regulations for supervised institutions and has since been folded into the current Manual of Regulations for Banks. Circular 706 established that a monitoring system has to exist. M-2023-013 is the more recent, more specific standard on what a well-run one demonstrably does, and it's the anchor a calibration programme should be built against, not the older circular alone.
In practice, a defensible calibration record for a single rule holds six things. The threshold value before the change, and the value after. The date of the change, and who approved it. The above-the-line sample and its productivity rate. The below-the-line sample and what it found. And the date of the next scheduled review. An institution that can produce this for its top ten rules on request is answering the evidence question. One that can only describe its process in general terms is not, whatever the process actually looks like day to day.
Steps to calibrate a transaction monitoring threshold
Pull the baseline. Before changing anything, capture the rule's current alert volume, productive-alert rate, and time-to-disposition over a fixed period, typically 30 to 90 days.
Run above-the-line testing. Sample a statistically reasonable slice of the alerts generated in that window and classify each as productive (escalated, filed, or otherwise actionable) or not. A productive rate that's very low signals the threshold is too loose.
Run below-the-line testing. Sample transactions that fell just under the threshold in the same window and review them against the typology the rule targets. Any that resemble genuine risk signal the threshold is too tight.
Model the proposed change against history. Before moving the threshold, test the candidate value against the same historical window and project the resulting alert volume and productivity rate. This is the pre-deployment testing step BSP's guidance expects, and it's what turns a threshold change from a guess into a measured one.
Document the decision. Record the before and after values, the test results from steps 2 and 3, who approved the change, and why. This record is the artifact an examiner asks for, not the fact that a meeting happened.
Schedule the next review. A threshold that was right in January is not guaranteed to be right in July, as customer mix, products, and typologies shift. Put a recurring review date on the calendar rather than waiting for the queue to complain again.
Repeat below-the-line testing even when the queue looks fine. A quiet queue with no below-the-line check is the most common way real risk goes undetected. It's evidence of nothing on its own.
Where does AI-assisted tuning fit into this process?
Carefully, and not yet as a replacement for the steps above. Pattern-detection tooling that watches alert trends over time and flags a rule that looks miscalibrated, before a human notices the queue has drifted, is a genuinely useful layer on top of a calibration programme. It shortens the time between a threshold going stale and someone catching it.
What it does not do, at least not without a person confirming the result, is decide the new threshold value or file the documentation an examiner will ask for. The above-the-line and below-the-line testing steps in this guide are analyst-led by design: a person samples the alerts, a person reviews the near-misses, a person signs the change. FyscalTech's approach to this treats AI-assisted tuning as a way to flag where to look, not a way to skip looking.
Any AI-generated recommendation still needs confirmation on what's shipping and how it should be described before specific claims are made about it externally. The honest framing today is that the platform supports analyst-led tuning, with pattern-detection assistance layered on top rather than replacing it.
How often should thresholds be reviewed?
On a fixed, documented schedule, not only in response to complaints about queue volume. A quarterly review of the top-volume rules, with above-the-line and below-the-line testing run each time, is a common baseline. That cadence should tighten around any material shift: a new product line, a new customer segment, or a typology an investigator has started seeing that the current rule set wasn't built to catch.
Where does threshold calibration fit in your AML program?
It's a Transaction Monitoring discipline with an AI Layer assist, not a standalone control. Calibration only works when the compliance team, not engineering, can adjust a threshold and re-test it, and when the platform can run a candidate threshold against historical data before it goes live. That's the same aggregation-and-dry-run capability our transaction monitoring pillar guide covers at a summary level; this piece is the deeper mechanics of the testing itself. Once a calibration decision produces a genuine alert, it flows into case management and, where confirmed, an STR filing under the same next-working-day standard as any other monitoring-driven suspicion, the same deadline institutions most often confuse with the CTR filing window. The filing itself is then a regulatory reporting problem, not a calibration one.

