A financial crime program is not one system. It is a chain: customer risk rating, sanctions and watchlist screening, transaction monitoring, alert triage, case investigation, and the filing itself. Each link rests on a different regulatory basis, tolerates a different kind of error, and gives a different answer to the only question that matters when a machine is inserted into it — how much of this decision may the machine make alone. Treating the chain as one AI project is how programs get into trouble, because the evidence an examiner wants at the screening layer is not the evidence an examiner wants at the disposition layer.
Two things moved in 2026 that change how that evidence has to be built. Both matter more than any modeling technique a program will choose this year.

The two changes that reset the ground
On April 17, 2026, the Federal Reserve, the OCC and the FDIC issued revised supervisory guidance on model risk management. The Federal Reserve carries it as SR 26-2; the OCC as Bulletin 2026-13. It "supersedes and replaces SR letter 11-7, Guidance on Model Risk Management (issued April 4, 2011) and SR letter 21-8, Interagency Statement on Model Risk Management for Bank Systems Supporting Bank Secrecy Act/Anti-Money Laundering Compliance (issued April 9, 2021)." Both documents are gone. The one interagency statement written specifically about AML systems has been folded back into general model risk guidance, and there is no BSA/AML-specific replacement.
On April 7, 2026, FinCEN proposed a rewrite of the AML/CFT program requirements themselves, published in the Federal Register on April 10 with comments closing June 9. The FDIC, OCC and NCUA issued conforming proposals the same day. The Federal Reserve published its own conforming proposal on July 9, 2026, with a sixty-day comment window, so that window is still open as this is written. The proposals would replace the long-standing "reasonably designed" formulation with a two-pronged test of whether a bank has established and maintained an effective, risk-based program, and would put a risk assessment process into the rule text — one that has to account for the national AML/CFT Priorities FinCEN publishes. FinCEN has proposed an effective date twelve months after a final rule issues.
Read together the direction is consistent: less prescription about method, more weight on whether the program produces useful output. Good news for teams who can measure their own systems, bad news for teams whose evidence is a binder describing a process.
The definition decides whether you have a model at all
The revised guidance narrows what counts. The term "model" now "refers to a complex quantitative method, system, or approach that applies statistical, economic, or financial theories to process input data into quantitative estimates." Then the sentence that changes AML architecture conversations: it "excludes simple arithmetic calculations, such as those found within spreadsheets, as well as deterministic rule-based processes and software where there are no statistical, economic, or financial theories underpinning their design or use."
A classic scenario-and-threshold monitoring engine is a deterministic rule-based process. On the face of the definition, it is not a model. The statistical work wrapped around it — above-the-line and below-the-line threshold testing, segmentation analysis, any layer that scores alerts — plainly is, as is a customer risk-rating model. A program validating its rules engine as a model is doing work the guidance no longer describes. A program treating its ML triage layer as "just a filter" is skipping work the guidance clearly describes.
There is a second scope line. A footnote states that generative AI and agentic AI models "are novel and rapidly evolving. As such, they are not within the scope of this guidance," while "the principles described in this guidance apply to traditional statistical and quantitative models and non-generative, non-agentic AI models." The agencies did not say generative systems are unregulated. They said this document does not govern them, and that an institution's own practices "should guide the determination of appropriate governance and controls for any tools, processes, or systems not covered in this document." Put a language model near a filing decision and you write that control framework yourself, and defend it, with no supervisory template to point at.
| Component of the stack | Statistical, economic or financial theory underneath? | Inside the April 2026 model definition | What still governs it |
|---|---|---|---|
| Scenario and threshold monitoring engine | No — deterministic rules | Outside, on the stated definition | The AML program rule. Tuning still has to be supported. |
| Threshold tuning and segmentation analysis | Yes | Inside | Conceptual soundness, outcomes analysis, monitoring |
| Alert scoring or triage model | Yes | Inside | Full lifecycle, plus suppression evidence |
| Customer risk rating model | Yes | Inside, and high materiality | Same, with heavier governance for regulatory purpose |
| Fuzzy name matching for sanctions screening | Depends on the algorithm | Read the implementation, not the label | Scoring-based matchers behave like models; string rules do not |
| Language model drafting narratives or case summaries | Generative | Expressly out of scope | Controls the institution designs and defends itself |
One boundary is worth knowing before a validation budget is set. The guidance "is expected to be most relevant to banking organizations with over $30 billion in total assets," and says excluding organizations at or below that line "is consistent with a tailored supervisory approach," while noting it may still be relevant to smaller institutions carrying significant model risk. The same $30 billion line marks the OCC's community bank minimum BSA/AML examination procedures, effective for examinations beginning February 1, 2026.
The footnote that changes the conversation
Buried at the end of the introduction: "This guidance does not set forth enforceable standards or prescriptive requirements; accordingly, non-compliance with this guidance will not result in supervisory criticism against a banking organization." Its own footnote adds the limit — "supervisory action may result for any violations of law or unsafe or unsound practices stemming from insufficient management of model risk."
That is not permission to stop validating. It relocates the obligation. The model risk guidance is not the enforceable thing. The program rule is. Build to the rule, and use the guidance as the method for proving you met it. A validation finding is not, by itself, a violation. A monitoring system that materially fails to detect what it was built to detect is a program problem however the validation report reads.
What an in-scope model is still asked to show
The substance of the old framework survived the rewrite, in tighter language. Validation still turns on conceptual soundness — assessing and documenting model design, key modeling choices, assumptions, qualitative judgments and data selection, with critical analysis of the quality and extent of developmental evidence. It still turns on outcomes analysis, which "compares model outputs to corresponding real-world outcomes to assess model performance relative to model objectives and business use," and the guidance is explicit that "when a model's design relies substantially on expert judgment, quantitative outcomes analysis helps to evaluate the quality of that judgment." That sentence is aimed squarely at threshold sets chosen by experienced BSA officers and never tested since. And it still turns on ongoing monitoring against changes in products, exposures, activities, clients, data relevance or market conditions.
Effective challenge keeps its central place, now defined as "the critical analysis conducted by objective experts who evaluate model risk and effect appropriate changes throughout the model lifecycle," performed by people with the expertise to challenge, the independence to stay objective, and "the organizational standing and influence to effect any change." The third clause is the one that fails in practice. Plenty of AML validation functions can write a finding. Fewer can stop a deployment.
Two details are easy to miss. Validation "generally occurs prior to a model's first use," but the guidance allows for urgent business need on condition of compensating controls such as limits on model use or closer performance monitoring. And validation quality "depends on the rigor and effectiveness of the review rather than on organizational structure," which puts the weight on what reviewers actually did rather than where they sit.
Where automated disposition is explicitly permitted
The most useful document in this area is not new and is routinely misread. OCC Interpretive Letter 1166, dated September 27, 2019, answers a bank's request to automate the filing of structuring-related suspicious activity reports. The OCC concluded that automated generation of a SAR narrative is consistent with its SAR regulation, given the bank's representation that filed reports would contain all required elements from the SAR form instructions and applicable FinCEN guidance. It further concluded that its SAR and program regulations "permit a bank to file a Structuring SAR based solely on an alert under the conditions and limitations described in your request letter." The OCC declined the bank's separate request for a regulatory sandbox and any forbearance.
The conditions are the content of the letter, and they read as a design specification. Alerts leave the automated path for manual investigation when there is any non-structuring alert, when the customer has prior case or SAR history, or when other high-risk activity sits close to the potential structuring. After a report is filed, the account is monitored and any alert within 90 days triggers manual review of original and subsequent transactions. A law enforcement request after filing triggers manual investigation and an amended filing if needed. And the bank runs parallel sample testing — manually reviewing a subset of filed reports to see whether more should have been included, amending where it should, and feeding results back into the process.
The OCC's reasoning explains why this transfers to some detection tasks and not others: "automation of any SAR filing process raises the possibility that suspicious activities will go unreported due to the absence of a more involved review. Accordingly, automated filing of Structuring SARs is only permissible to the extent that it is supported by strong risk governance that remove higher risk transactions from the automated process." Structuring is described as lower risk and more readily automated because the underlying judgment is simple. The letter also names the failure conditions plainly: an automated process that is not "regularly overseen, evaluated and updated," a system that "materially fails to identify instances of structuring," or one that "results in alert backlogs" raises issues under both the SAR and program regulations.
The category line is inside the regulation, not inside the model
The SAR regulation itself separates the cases. For structuring, the filing trigger is that the institution "knows, suspects, or has reason to suspect" the transaction is designed to evade a Bank Secrecy Act requirement. For the no-apparent-purpose category, a report is required when the institution knows of no reasonable explanation for the transaction after examining the available facts, including the background and possible purposes of the transaction. One standard can be met by an objective read of transaction values and frequency. The other has an examination step written into it.
The engineering consequence is that alerts should be routed by regulatory basis first and model score second. A pipeline that ranks every alert on one probability and sends the top slice to analysts has erased the distinction the regulation makes. A pipeline that partitions on basis, then automates only inside the partition where the published record supports it, has an answer when someone asks why this alert closed unread.
How much published support exists for automating each step
Our editorial read of how much explicit regulatory support exists in the public record for automating each step. Not a measured statistic.
These figures read the published record; they measure nothing. The top rows have named support — an interpretive letter, an interagency FAQ, or regulatory text describing the judgment as objective. The bottom row has none. No published statement permits closing a no-apparent-purpose alert with no human examination of the facts, and a program doing it at volume should expect that question first.
The October 2025 SAR FAQs deleted work most programs were still doing
On October 9, 2025, FinCEN issued four FAQs jointly with the Federal Reserve, FDIC, NCUA and OCC, stating that the answers "do not alter existing BSA legal or regulatory requirements or establish new supervisory expectations." Each removes work and each has a system consequence.
Threshold proximity is not, by itself, a reason to file. "The mere presence of a transaction or series of transactions by or on behalf of the same person at or near the $10,000 CTR threshold is not information sufficient to require the filing of a SAR." Filing is required where the institution knows, suspects or has reason to suspect the activity is designed to evade the reporting requirement. Any scenario that fires on proximity alone is producing alerts that carry no filing obligation, and the alerts still have to be worked.
The 90-day continuing-activity review is not required. The FAQs trace the practice to a suggestion FinCEN made in October 2000 that hardened into a perceived requirement. The answer is that an institution "is not required to conduct a separate review — manual or otherwise — of a customer or account following the filing of a SAR to determine whether suspicious activity has continued," and may instead rely on risk-based internal policies, procedures and controls. For institutions that do elect to follow the continuing-activity guidance, the FAQs lay out the timeline: detection at day 0, initial filing by day 30, end of the 90-day period at day 120, continuing-activity filing at day 150.
There is no requirement to document a decision not to file. Where an institution chooses to, "a short, concise statement documenting a financial institution's SAR decision will likely suffice," with more only in complex scenarios. Programs that built a mandatory no-file justification form and trained analysts to fill it out at length created an artifact no rule asked for — one that becomes discovery material about every alert the model suppressed.
Scale matters here. Published analysis of FinCEN's SAR Stats puts 2025 filings by banks, savings associations and credit unions above 2.19 million, and all filer categories above 4.1 million. Removing one avoidable review per filing is a larger capacity change than most model improvements deliver.
Explainability that survives an examination
"Explainable" in a BSA/AML setting does not mean a feature attribution chart. It means an examiner picks an alert from eighteen months ago and the institution reproduces the decision: model version, feature values as of that timestamp, the threshold in force that day, the routing rule, the person who dispositioned it, and the regulatory basis for closure. If any of that came from a table since overwritten, the decision is not reproducible and the explanation is a reconstruction.
That requirement drives architecture more than model choice does. Point-in-time feature storage rather than recomputation from current data. Model and threshold versions carried on the alert record, not looked up later. Closure reasons stored as codes tied to the regulatory category, not as free text. Systems built this way answer in an afternoon. The rest rebuild history under time pressure, and the rebuild is what draws the finding.
Suppression is the other half. Any model that reduces analyst workload does so by deciding some alerts do not need a human, and the evidence for that is a different artifact from a precision score. We have written separately on how to build alert triage so the suppression answer exists before anyone asks for it, including why historical dispositions make a contaminated training label.
Screening is a different problem from monitoring
Sanctions and watchlist screening looks adjacent to transaction monitoring and behaves nothing like it. Monitoring tolerates a missed alert as a tuning question. Screening does not, because sanctions liability does not depend on the institution having intended anything. That makes screening a recall-first problem, where tuning is entirely about how much review capacity the institution will spend to keep recall high.
It is also not a list lookup, for a structural reason. OFAC's 50 Percent Rule treats an entity owned fifty percent or more, in the aggregate, by one or more blocked persons as blocked itself, whether or not it appears on any list. Matching a payment party against published names does not answer the ownership question. Answering it needs ownership data the institution does not own, resolved across sources that spell the same company four ways — an entity resolution problem before it is a screening problem. Treat screening as string matching against a downloaded list and you have solved the visible half.
Vendor models and the part of the contract nobody negotiated
Most detection stacks are bought, not built. The revised guidance devotes its final section to that and does not soften the expectation: customized vendor products "can present unique challenges for validation," and because components may be proprietary, an institution "may not receive from the vendor the underlying code, data, or methodology that they would have if a model were developed internally. Nevertheless, the principles of model risk management remain applicable."
Sound practice, per the guidance, includes "developing an understanding of the vendor model, including its conceptual soundness, design, development data, and performance," ongoing monitoring to assess whether vendor models "are accurate, remain fit for purpose, and continue to be reliable," and evaluating any customization as part of validation. Every one of those is a contract term, and the time to secure it is before signature. Ask for developmental evidence after go-live and you get a marketing deck. Make documentation delivery, benchmark data and a right to independent testing part of the purchase and the evidence exists, because someone was paid to produce it.
A build sequence that produces the evidence
Order of work for a detection change that must survive examination
The sequence is front-loaded on data and governance for a reason. Programs that get the model working first tend to discover, late, that they cannot reconstruct why any individual alert closed the way it did. Reversing that order is expensive, and the cost lands during an examination rather than during a sprint.
The evidence pack, in the order someone will ask for it
- Inventory entry per in-scope model, with materiality reasoning stated
- Conceptual soundness memo: design, assumptions, data selection, and the choices rejected
- Outcomes analysis against real dispositions, with the performance thresholds set in advance
- Suppression analysis on a holdout the model never saw, sized to detect what you claim to detect
- Effective challenge record showing what the challenger changed, not only what they reviewed
- Point-in-time reproduction of a sample of closed alerts, run cold
- Vendor documentation, customization log, and independent testing rights on file
- Ongoing monitoring reports with the trigger that would force recalibration written down
What is genuinely unsettled
Three things are in motion. A program planning past this year should treat them as open, not decided.
The program rule is not final. FinCEN's proposal and the banking agencies' conforming proposals are proposals, and the Federal Reserve's comment window is still open. If a final rule issues in the proposed form, the effective date falls twelve months later — enough time to rebuild a risk assessment process, not enough to rebuild a data platform.
Generative and agentic systems sit outside the model guidance. The agencies excluded them and told institutions to determine appropriate governance themselves. That is genuine ambiguity, not a loophole. Use a language model to draft filing narratives and you must be ready to explain, with no supervisory template, how you keep a fabricated fact out of a report sent to law enforcement.
The reporting thresholds may move. The Financial Reporting Threshold Modernization Act, H.R. 1799, would raise the currency transaction reporting threshold from $10,000 to $30,000 and index it periodically. It advanced out of the House Financial Services Committee in January 2026 and has not been enacted, so nothing should be built on the assumption that it will be. But a stack with thresholds hard-coded into scenario logic rather than held as configuration will absorb a real project if the number changes.
Bottom line
The 2026 changes reward programs that can measure themselves and expose programs that describe themselves. The model definition narrowed, so validation effort should move off rules engines and onto the statistical layers that were getting a lighter touch. The guidance stopped claiming to be enforceable, which puts the weight back on the program rule and on outcomes. And the parts of the public record that authorize automation — an interpretive letter about structuring, an FAQ about continuing activity — authorize it narrowly, on conditions, with an audit loop attached. Build inside those conditions and the evidence exists when someone asks.
Frequently asked questions
No. SR 26-2, issued April 17, 2026 by the Federal Reserve, OCC and FDIC, supersedes and replaces both SR 11-7 and SR 21-8, the 2021 interagency statement on model risk management for BSA/AML systems. The substance of validation carried over, but the scope, the model definition and the enforceability language all changed.
Under the April 2026 definition, deterministic rule-based processes with no statistical, economic or financial theory behind their design are excluded from the term. The statistical work around the engine — threshold tuning, segmentation, any scoring layer — is not excluded. The rules engine is still governed by the AML program rule regardless.
The OCC concluded in Interpretive Letter 1166 that a bank may file a structuring-related report based solely on an alert, under specific conditions: guardrails that pull higher-risk alerts out for manual investigation, monitoring and manual review if any alert recurs within 90 days, manual handling of law enforcement follow-ups, and post-filing sample testing that feeds corrections back into the process. That reasoning is specific to structuring, which the letter describes as involving simpler judgment.
The October 9, 2025 interagency FAQs state that an institution is not required to conduct a separate review, manual or otherwise, following a filing to determine whether suspicious activity has continued, and may rely on risk-based internal policies, procedures and controls instead. Institutions that elect to follow the continuing-activity guidance have a stated timeline running to a day-150 filing.
In practice, the ability to reproduce a specific past decision: model version, feature values as of that timestamp, threshold in force, routing rule, disposition, and the regulatory basis for closure. Feature attribution helps a validator understand the model. Point-in-time reproduction is what answers the examination question.