ML Pipeline for Financial Monitoring: Why Classification Alone Isn't Enough
Published: 2026-09-27 · Author: AI Release · @ai_release1
⚡ The Gist in 5 Seconds - Senior data engineer Maxim Kotsyuba (experience at VTB and Sber) built an end-to-end ML/MLOps pipeline for detecting suspicious transactions and brought it to working code. - Instead of binary classification, the model ranks operations by risk level: this lets analysts work with a prioritized queue rather than drowning in false positives. - The project's source code and MVP are published on GitHub — but this isn't a ready-made guide; it's an analysis of where the standard approach falls short. ### 🔍 What Was Found Maxim Kotsyuba, a senior data engineer with experience at VTB and Sber and a graduate of the online master's program in Data Engineering from HSE University and Netology, built an end-to-end ML/MLOps pipeline for his capstone project. It doesn't just classify transactions — it ranks them by risk level. The problem is real: existing rule-based systems produce over 90% false positives, and reviewing this flood consumes most of analysts' working time. The workload has grown manifold over 15 years: from 2010 to 2025, card transaction volume increased from 3.1 to 72.7 billion (23x), while the number of operating credit organizations dropped from 1,058 to 353 (nearly 67%). Per organization, the load grew roughly 70-fold — from 2.9 to 206 million transactions per entity. The experiments used the SAML-D dataset, where the share of suspicious transactions is just 0.119%. With such severe imbalance, ranking metrics like Precision@k and Lift@k proved far more informative than classification metrics. For the top 100, both models achieved Precision@100 = 1.00. For the top 500, the figures naturally declined: Precision@500 was 0.322 for LightGBM and 0.310 for XGBoost. The author calls the 0.5 classification threshold simply unfit for production: in reality, it should be determined by the acceptable analyst workload and the required detection recall. ### 💡 Why It Matters The key takeaway: the model is useful not as a final decision-making tool, but as a risk-based ranking mechanism. It produces a risk score on the basis of which transactions are ordered by potential risk, and the analyst first handles the riskiest ones. This solves the bottleneck of manual review as transaction volumes grow. Moreover, sanctions increase the need for locally deployable solutions: foreign software and infrastructure may lose availability, licensing, and technical support, so it's important to have a reproducible setup that can be maintained independently. ### 🧩 Context When framing the task, rule-based mechanisms, ready-made commercial AML platforms (SAS AML, NICE Actimize), and in-house development by the bank itself were all evaluated. Each approach had limitations: rule-based systems take time to adapt to new fraud schemes, commercial platforms remain closed black boxes where the logic is hard to examine, and internal development runs into a manual, non-automated model training process. Rules in financial monitoring are still necessary due to their transparency, but
⚡ The Gist in 5 Seconds - Senior data engineer Maxim Kotsyuba (experience at VTB and Sber) built an end-to-end ML/MLOps pipeline for detecting suspicious transactions and brought it to working code.
- Instead of binary classification, the model ranks operations by risk level: this lets analysts work with a prioritized queue rather than drowning in false positives.
- The project's source code and MVP are published on GitHub — but this isn't a ready-made guide; it's an analysis of where the standard approach falls short.
🔍 What Was Found Maxim Kotsyuba, a senior data engineer with experience at VTB and Sber and a graduate of the online master's program in Data Engineering from HSE University and Netology, built an end-to-end ML/MLOps pipeline for his capstone project.
It doesn't just classify transactions — it ranks them by risk level.
The problem is real: existing rule-based systems produce over 90% false positives, and reviewing this flood consumes most of analysts' working time.
The workload has grown manifold over 15 years: from 2010 to 2025, card transaction volume increased from 3.1 to 72.7 billion (23x), while the number of operating credit organizations dropped from 1,058 to 353 (nearly 67%).
Per organization, the load grew roughly 70-fold — from 2.9 to 206 million transactions per entity.