Reported gains from graph-based machine learning in anti-money laundering and fraud detection can vary substantially with the historical evidence available to a model and with the evaluation protocol used. This study introduces a coverage-aware and protocol-aware evaluation framework for assessing account-risk enrichment, graph neural networks, tabular gradient boosting, and graph-tabular stacking across synthetic banking anti-money-laundering data, a real Bitcoin transaction graph, and a complementary real credit-card fraud dataset. The analysis compares account-risk enrichment, four topology-only graph neural networks, XGBoost, LightGBM, and embedding-based stacking under warm-start, account-grouped, cold-start, coverage-sensitivity, and temporal evaluation settings. On the synthetic HI-Small benchmark, strict account grouping removes source-account historical coverage by construction and substantially limits the opportunity for enrichment to contribute useful signal. Under the warm-start protocol, enrichment produces only a small change in discrimination, while the coverage-sensitivity analysis shows no reliable improvement in ROC-AUC as historical coverage increases. PR-AUC results instead indicate a small but consistent performance cost at most tested coverage levels after correction for multiple comparisons. On the lower-coverage LI-Small benchmark, enrichment again shows a small negative effect that does not remain significant after manuscript-wide correction. External validation on the Elliptic Bitcoin graph shows that topology-only graph neural networks do not automatically outperform strong tabular models. GraphSAGE is the strongest graph neural network tested, but XGBoost and LightGBM achieve clearly higher ROC-AUC and PR-AUC. Graph-tabular stacking is also dataset dependent: it yields small positive gains on Elliptic, while tabular features alone are competitive with or superior to stacking on HI-Small. These results show that graph-based improvements should not be interpreted independently of historical coverage, split construction, and feature redundancy. Practical evaluation should report coverage and protocol alongside graph-derived claims, use strong tabular baselines, and complement ROC-AUC with PR-AUC and calibration-oriented metrics before concluding that a graph-based method provides a robust advantage.
| Published in | International Journal of Data Science and Analysis (Volume 12, Issue 5) |
| DOI | 10.11648/j.ijdsa.20261205.11 |
| Page(s) | 97-113 |
| Creative Commons |
This is an Open Access article, distributed under the terms of the Creative Commons Attribution 4.0 International License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution and reproduction in any medium or format, provided the original work is properly cited. |
| Copyright |
Copyright © The Author(s), 2026. Published by Science Publishing Group |
Anti-money Laundering, Graph Neural Networks, Protocol-aware Evaluation, Coverage-aware Evaluation, Account-risk Enrichment, Multi-dataset Validation, Stacking
| [1] | T. N. Kipf and M. Welling, "Semi-Supervised Classification with Graph Convolutional Networks," in Proceedings of the International Conference on Learning Representations (ICLR), 2017. |
| [2] | W. L. Hamilton, Z. Ying, and J. Leskovec, "Inductive Representation Learning on Large Graphs," in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017. |
| [3] | P. Velickovic, G. Cucurull, A. Casanova, A. Romero, P. Liò, and Y. Bengio, "Graph Attention Networks," in Proceedings of the International Conference on Learning Representations (ICLR), 2018. |
| [4] | K. Xu, W. Hu, J. Leskovec, and S. Jegelka, "How Powerful Are Graph Neural Networks?" in Proceedings of the International Conference on Learning Representations (ICLR), 2019. |
| [5] | T. Chen and C. Guestrin, "XGBoost: A Scalable Tree Boosting System," in Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), 2016, pp. 785-794. |
| [6] | G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T.-Y. Liu, "LightGBM: A Highly Efficient Gradient Boosting Decision Tree," in Advances in Neural Information Processing Systems (NeurIPS), vol. 30, 2017. |
| [7] | Y. Liu, X. Ao, Z. Qin, J. Chi, J. Feng, H. Yang, and Q. He, "Pick and Choose: A GNN-Based Imbalanced Learning Approach for Fraud Detection," in Proceedings of the Web Conference (WWW), 2021, pp. 3168-3177. |
| [8] | Y. Dou, Z. Liu, L. Sun, Y. Deng, H. Peng, and P. S. Yu, "Enhancing Graph Neural Network-based Fraud Detectors against Camouflaged Fraudsters," in Proceedings of the ACM International Conference on Information and Knowledge Management (CIKM), 2020, pp. 315-324. |
| [9] | M. Weber, G. Domeniconi, J. Chen, D. K. I. Weidele, C. Bellei, T. Robinson, and C. E. Leiserson, "Anti-Money Laundering in Bitcoin: Experimenting with Graph Convolutional Networks for Financial Forensics," in Proceedings of the KDD Workshop on Anomaly Detection in Finance, 2019. |
| [10] | E. Altman, J. Blanua, L. von Niederhäusern, B. Egressy, A. Anghel, and K. Atasu, "Realistic Synthetic Financial Transactions for Anti-Money Laundering Models," in Proceedings of the NeurIPS Datasets and Benchmarks Track, 2023. |
| [11] | B. Johannessen and M. Jullum, "Finding Money Launderers Using Heterogeneous Graph Neural Networks," arXiv preprint, 2023. |
| [12] | L. Akoglu, H. Tong, and D. Koutra, "Graph-Based Anomaly Detection and Description: A Survey," Data Mining and Knowledge Discovery, vol. 29, no. 3, pp. 626-688, 2015. |
| [13] | V. Van Vlasselaer, C. Bravo, O. Caelen, T. Eliassi-Rad, L. Akoglu, M. Snoeck, and B. Baesens, "APATE: A Novel Approach for Automated Credit Card Transaction Fraud Detection Using Network-Based Extensions," Decision Support Systems, vol. 75, pp. 38-48, 2015. |
| [14] | T. Pourhabibi, K.-L. Ong, B. H. Kam, and Y. L. Boo, "Fraud Detection: A Systematic Literature Review of Graph-Based Anomaly Detection Approaches," Decision Support Systems, vol. 133, pp. 113303, 2020. |
| [15] | S. Kaufman, S. Rosset, C. Perlich, and O. Stitelman, "Leakage in Data Mining: Formulation, Detection, and Avoidance," ACM Transactions on Knowledge Discovery from Data (TKDD), vol. 6, no. 4, pp. 1-21, 2012. |
| [16] | D. Micci-Barreca, "A Preprocessing Scheme for High-Cardinality Categorical Attributes in Classification and Prediction Problems," ACM SIGKDD Explorations Newsletter, vol. 3, no. 1, pp. 27-32, 2001. |
| [17] | L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin, "CatBoost: Unbiased Boosting with Categorical Features," in Advances in Neural Information Processing Systems (NeurIPS), vol. 31, 2018. |
| [18] | A. Dal Pozzolo, O. Caelen, R. A. Johnson, and G. Bontempi, "Calibrating Probability with Undersampling for Unbalanced Classification," in Proceedings of the 2015 IEEE Symposium Series on Computational Intelligence (SSCI), Cape Town, South Africa, 2015, pp. 159-166. |
| [19] | C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, "On Calibration of Modern Neural Networks," in Proceedings of the 34th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, vol. 70, 2017, pp. 1321-1330. |
| [20] | S. Maganti, "When Graph Structure Becomes a Liability: A Critical Re-Evaluation of Graph Neural Networks for Bitcoin Fraud Detection under Temporal Distribution Shift," arXiv preprint, 2026. |
| [21] | K. Hayat and B. Magnier, "Data Leakage and Deceptive Performance: A Critical Examination of Credit Card Fraud Detection Methodologies," Mathematics, vol. 13, no. 16, 2563, 2025. |
| [22] | H. Khaleghpour and B. McKinney, "Leakage Safe Graph Features for Interpretable Fraud Detection in Temporal Transaction Networks," arXiv preprint, 2026. |
| [23] | B. Deprez, W. Wei, W. Verbeke, B. Baesens, K. Mets, and T. Verdonck, "Advances in Continual Graph Learning for Anti-Money Laundering Systems: A Comprehensive Review," WIREs Computational Statistics, vol. 17, no. 3, 2025. |
| [24] | W. Zhang, R. Li, Q. Bai, and S. Wang, "SparseFraudNet: A Graph-based Approach for Cold-start Fraud Detection with Information Aggregation," ACM Transactions on Information Systems, 2025. |
| [25] | S. Seabold and J. Perktold, "statsmodels: Econometric and Statistical Modeling with Python," in Proceedings of the 9th Python in Science Conference (SciPy), 2010. |
APA Style
Bettaieb, K., Echi, A. K., Hammami, H. (2026). Coverage and Protocol-Aware Evaluation of Risk Enrichment and Graph-Tabular Learning for Anti-Money Laundering and Fraud Detection. International Journal of Data Science and Analysis, 12(5), 97-113. https://doi.org/10.11648/j.ijdsa.20261205.11
ACS Style
Bettaieb, K.; Echi, A. K.; Hammami, H. Coverage and Protocol-Aware Evaluation of Risk Enrichment and Graph-Tabular Learning for Anti-Money Laundering and Fraud Detection. Int. J. Data Sci. Anal. 2026, 12(5), 97-113. doi: 10.11648/j.ijdsa.20261205.11
AMA Style
Bettaieb K, Echi AK, Hammami H. Coverage and Protocol-Aware Evaluation of Risk Enrichment and Graph-Tabular Learning for Anti-Money Laundering and Fraud Detection. Int J Data Sci Anal. 2026;12(5):97-113. doi: 10.11648/j.ijdsa.20261205.11
@article{10.11648/j.ijdsa.20261205.11,
author = {Karim Bettaieb and Afef Kacem Echi and Houcem Hammami},
title = {Coverage and Protocol-Aware Evaluation of Risk Enrichment and Graph-Tabular Learning for Anti-Money Laundering and Fraud Detection},
journal = {International Journal of Data Science and Analysis},
volume = {12},
number = {5},
pages = {97-113},
doi = {10.11648/j.ijdsa.20261205.11},
url = {https://doi.org/10.11648/j.ijdsa.20261205.11},
eprint = {https://article.sciencepublishinggroup.com/pdf/10.11648.j.ijdsa.20261205.11},
abstract = {Reported gains from graph-based machine learning in anti-money laundering and fraud detection can vary substantially with the historical evidence available to a model and with the evaluation protocol used. This study introduces a coverage-aware and protocol-aware evaluation framework for assessing account-risk enrichment, graph neural networks, tabular gradient boosting, and graph-tabular stacking across synthetic banking anti-money-laundering data, a real Bitcoin transaction graph, and a complementary real credit-card fraud dataset. The analysis compares account-risk enrichment, four topology-only graph neural networks, XGBoost, LightGBM, and embedding-based stacking under warm-start, account-grouped, cold-start, coverage-sensitivity, and temporal evaluation settings. On the synthetic HI-Small benchmark, strict account grouping removes source-account historical coverage by construction and substantially limits the opportunity for enrichment to contribute useful signal. Under the warm-start protocol, enrichment produces only a small change in discrimination, while the coverage-sensitivity analysis shows no reliable improvement in ROC-AUC as historical coverage increases. PR-AUC results instead indicate a small but consistent performance cost at most tested coverage levels after correction for multiple comparisons. On the lower-coverage LI-Small benchmark, enrichment again shows a small negative effect that does not remain significant after manuscript-wide correction. External validation on the Elliptic Bitcoin graph shows that topology-only graph neural networks do not automatically outperform strong tabular models. GraphSAGE is the strongest graph neural network tested, but XGBoost and LightGBM achieve clearly higher ROC-AUC and PR-AUC. Graph-tabular stacking is also dataset dependent: it yields small positive gains on Elliptic, while tabular features alone are competitive with or superior to stacking on HI-Small. These results show that graph-based improvements should not be interpreted independently of historical coverage, split construction, and feature redundancy. Practical evaluation should report coverage and protocol alongside graph-derived claims, use strong tabular baselines, and complement ROC-AUC with PR-AUC and calibration-oriented metrics before concluding that a graph-based method provides a robust advantage.},
year = {2026}
}
TY - JOUR T1 - Coverage and Protocol-Aware Evaluation of Risk Enrichment and Graph-Tabular Learning for Anti-Money Laundering and Fraud Detection AU - Karim Bettaieb AU - Afef Kacem Echi AU - Houcem Hammami Y1 - 2026/09/08 PY - 2026 N1 - https://doi.org/10.11648/j.ijdsa.20261205.11 DO - 10.11648/j.ijdsa.20261205.11 T2 - International Journal of Data Science and Analysis JF - International Journal of Data Science and Analysis JO - International Journal of Data Science and Analysis SP - 97 EP - 113 PB - Science Publishing Group SN - 2575-1891 UR - https://doi.org/10.11648/j.ijdsa.20261205.11 AB - Reported gains from graph-based machine learning in anti-money laundering and fraud detection can vary substantially with the historical evidence available to a model and with the evaluation protocol used. This study introduces a coverage-aware and protocol-aware evaluation framework for assessing account-risk enrichment, graph neural networks, tabular gradient boosting, and graph-tabular stacking across synthetic banking anti-money-laundering data, a real Bitcoin transaction graph, and a complementary real credit-card fraud dataset. The analysis compares account-risk enrichment, four topology-only graph neural networks, XGBoost, LightGBM, and embedding-based stacking under warm-start, account-grouped, cold-start, coverage-sensitivity, and temporal evaluation settings. On the synthetic HI-Small benchmark, strict account grouping removes source-account historical coverage by construction and substantially limits the opportunity for enrichment to contribute useful signal. Under the warm-start protocol, enrichment produces only a small change in discrimination, while the coverage-sensitivity analysis shows no reliable improvement in ROC-AUC as historical coverage increases. PR-AUC results instead indicate a small but consistent performance cost at most tested coverage levels after correction for multiple comparisons. On the lower-coverage LI-Small benchmark, enrichment again shows a small negative effect that does not remain significant after manuscript-wide correction. External validation on the Elliptic Bitcoin graph shows that topology-only graph neural networks do not automatically outperform strong tabular models. GraphSAGE is the strongest graph neural network tested, but XGBoost and LightGBM achieve clearly higher ROC-AUC and PR-AUC. Graph-tabular stacking is also dataset dependent: it yields small positive gains on Elliptic, while tabular features alone are competitive with or superior to stacking on HI-Small. These results show that graph-based improvements should not be interpreted independently of historical coverage, split construction, and feature redundancy. Practical evaluation should report coverage and protocol alongside graph-derived claims, use strong tabular baselines, and complement ROC-AUC with PR-AUC and calibration-oriented metrics before concluding that a graph-based method provides a robust advantage. VL - 12 IS - 5 ER -