Neuro-Behavioral Fusion for Predictive Tax Fraud Detection in Cloud-Based Government Systems
DOI:
https://doi.org/10.5281/zenodo.20742667Keywords:
Cloud computing, behavioural analytics, detect, fraud, government systems, hybrid architectures, hybrid deep learning, predictive, tax, Twitter.Abstract
Accurate prediction of tax fraud fosters favorable taxpayer-business relationships, streamlines tax authorities' operations, and optimally directs government investments to enhance public services. Reliable prediction of fraudulent taxpayers, however, requires careful selection of assessed features because of pronounced privacy, ethical, and security concerns associated with government-related data. Furthermore, limited historical records are available because fraud often remains undetected. Consequently, effective machine learning-based predictive models must deploy hybrid architectures capable of learning from different types of features, integrating feature engineering with representation learning, and extracting fraud behavioral patterns during training. Moreover, tax fraud detection is a pattern-discovery problem with critical imbalance between the positive and negative classes, necessitating the use of behaviorally informative indicators that optimize predictive performance, fairness, and robustness.
Predictive models should therefore integrate hybrid deep learning representations with behavioral-based tax fraud indicators for risk detection. Cloud-enabled government data ecosystems provide abundant data sources for detecting and predicting all forms of tax fraud, and during operation can be simulated to include as many behaviors that are indicators of tax fraud as possible. Predictive models can thus be trained and tested using a behavioral pattern-discovery approach, forming the foundation for hybrid deep learning-based models that use a dual-tower architecture to employ both behavioral indicators and independently engineered features to optimize detection risk.
References
[1] Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., … Zheng, X. (2016). TensorFlow: A system for large-scale machine learning. In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’16) (pp. 265–283). USENIX Association.
[2] Akoglu, L., Tong, H., & Koutra, D. (2015). Graph based anomaly detection and description: A survey. Data Mining and Knowledge Discovery, 29(3), 626–688.
[3] Arrieta, A. B., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115.
[4] Audet, A., & Neyshabur, B. (2023). Implicit bias of optimization algorithms in deep learning: A survey. Journal of Machine Learning Research, 24, 1–89.
[5] Bahdanau, D., Cho, K., & Bengio, Y. (2015). Neural machine translation by jointly learning to align and translate. In Proceedings of the 3rd International Conference on Learning Representations (ICLR).
[6] Bao, W., Yue, J., & Rao, Y. (2017). A deep learning framework for financial time series using stacked autoencoders and long-short term memory. PLOS ONE, 12(7), e0180944.
[7] Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115.
[8] Biecek, P. (2018). DALEX: Explainers for complex predictive models in R. Journal of Machine Learning Research, 19, 1–5.
[9] Bishop, C. M. (2006). Pattern recognition and machine learning. Springer.
[10] Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32.
[11] Budach, L., Feuerpfeil, J., Ihde, N., Nathansen, A., Noack, D., Patzlaff, N., Naumann, F., & Spiekermann, M. (2022). The effects of data augmentation on deep learning for tabular data. Proceedings of the VLDB Endowment, 15(8), 1580–1593.
[12] Carcillo, F., Le Borgne, Y. A., Caelen, O., Kessaci, Y., Oblé, F., & Bontempi, G. (2019). Combining unsupervised and supervised learning in credit card fraud detection. Information Sciences, 557, 317–331.
[13] Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), 1–58.
[14] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM.
[15] Christodoulou, E., Ma, J., Collins, G. S., Steyerberg, E. W., Verbakel, J. Y., & Van Calster, B. (2019). A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. Journal of Clinical Epidemiology, 110, 12–22.
[16] Cloudera. (2020). Modern data architecture and governance in the cloud. Cloudera White Paper.
[17] Cohen, W. W., Ravikumar, P., & Fienberg, S. E. (2003). A comparison of string distance metrics for name-matching tasks. In Proceedings of the IJCAI Workshop on Information Integration on the Web (pp. 73–78).
[18] Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297.
[19] Dehghani, M., Shakeri, H., & Zohrevand, A. (2021). Fraud detection in financial statements using deep learning and ensemble approaches: A systematic review. Expert Systems with Applications, 168, 114–168.
[20] Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT 2019 (pp. 4171–4186). Association for Computational Linguistics.
[21] Doersch, C. (2016). Tutorial on variational autoencoders. arXiv preprint arXiv:1606.05908.
[22] Dua, D., & Graff, C. (2019). UCI machine learning repository. University of California, Irvine.
[23] Esteva, A., Robicquet, A., Ramsundar, B., Kuleshov, V., DePristo, M., Chou, K., Cui, C., Corrado, G., Thrun, S., & Dean, J. (2019). A guide to deep learning in healthcare. Nature Medicine, 25(1), 24–29.
[24] Feurer, M., Klein, A., Eggensperger, K., Springenberg, J. T., Blum, M., & Hutter, F. (2015). Efficient and robust automated machine learning. In Advances in Neural Information Processing Systems (pp. 2962–2970).
[26] Friedler, S. A., Scheidegger, C., Venkatasubramanian, S., Choudhary, S., Hamilton, E. P., & Roth, D. (2019). A comparative study of fairness-enhancing interventions in machine learning. In Proceedings of FAT* 2019 (pp. 329–338). ACM.
[27] Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), 1–37.
[26] Garfinkel, S. L. (2015). De-identification of personal information. NIST Internal Report.
[27] Ghorbani, A., Wexler, J., Zou, J., & Kim, B. (2019). Towards automatic concept-based explanations. In Advances in Neural Information Processing Systems (pp. 9277–9286).
[28] Gilpin, L. H., Bau, D., Yuan, B. Z., Bajwa, A., Specter, M., & Kagal, L. (2018). Explaining explanations: An overview of interpretability of machine learning. In 2018 IEEE 5th International Conference on Data Science and Advanced Analytics (DSAA) (pp. 80–89). IEEE.
[29] Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.
[30] Gopalan, P. K., Hofman, J. M., & Blei, D. M. (2012). Scalable recommendation with Poisson factorization. In Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence (pp. 326–335).
[31] Guo, C., & Berkhahn, F. (2016). Entity embeddings of categorical variables. arXiv preprint arXiv:1604.06737.
[32] Hamilton, W. L. (2020). Graph representation learning. Morgan & Claypool.
[33] Han, J., Kamber, M., & Pei, J. (2011). Data mining: Concepts and techniques (3rd ed.). Morgan Kaufmann.
Additional Files
Published
Issue
Section
License
The American Online Journal of Science and Engineering (AOJSE) is an open-access journal, and all published articles are made freely available to readers worldwide under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0). This license governs the use, sharing, adaptation, and distribution of all content published in AOJSE and is designed to promote the broadest possible dissemination and reuse of scholarly research.
Creative Commons Attribution 4.0 International License (CC BY 4.0)
Under the Creative Commons Attribution 4.0 International License, any user is free to share, copy, and redistribute the published material in any medium or format, and to adapt, remix, transform, and build upon the material for any purpose, including commercial use, provided that appropriate credit is given to the original authors and source. The license cannot be revoked as long as the terms are followed.
When using or redistributing content published in AOJSE, users must provide proper attribution by citing the original authors' names, the article title, the journal name (AOJSE), the volume and issue number, the year of publication, and the DOI or URL of the article. Users must also indicate if any changes were made to the original work. Attribution must be provided in a reasonable manner but must not suggest that the authors or AOJSE endorse the user or their use of the material.
Author Rights and Copyright Retention
AOJSE respects and upholds the intellectual property rights of all contributing authors. Authors who publish in AOJSE retain full copyright over their work. By submitting a manuscript to AOJSE, authors grant the journal a non-exclusive, worldwide, royalty-free license to publish, reproduce, distribute, publicly display, and archive the article in print and electronic formats. This license allows AOJSE to make the article freely accessible to readers globally while ensuring that the authors remain the rightful owners of their work.
Authors are permitted to deposit their published articles in institutional repositories, personal websites, academic networking platforms, and preprint servers, provided that the original publication in AOJSE is acknowledged and properly cited. Authors may also reuse their published content in subsequent works, presentations, teaching materials, and grant applications without requiring prior permission from the journal.
Third-Party Content
Authors are solely responsible for obtaining written permission to reproduce any third-party material, including figures, tables, images, data, or excerpts from other publications, included in their manuscript. Evidence of such permissions must be provided to the editorial office upon request. AOJSE does not assume any responsibility for copyright infringement arising from the unauthorized inclusion of third-party content in published articles.
Permitted Uses
Under the CC BY 4.0 license, the following uses of AOJSE published content are freely permitted without prior written permission from the journal or authors, provided that proper attribution is given. Users may read, download, print, and distribute articles for personal, educational, or research purposes. Researchers may reuse data, figures, and findings for meta-analyses, systematic reviews, and secondary research. Educators may incorporate published articles into course materials, syllabi, and academic presentations. Journalists, policymakers, and practitioners may reference and cite published findings for professional and public interest purposes.
Prohibited Uses
While the CC BY 4.0 license is broad and permissive, users may not apply legal terms or technological measures that legally restrict others from doing anything the license permits. Users may not misrepresent the original authorship of published content or imply endorsement by the original authors or AOJSE without explicit consent. Plagiarism, data fabrication, and any form of academic misconduct in the reuse of published material are strictly prohibited and are a violation of publication ethics standards.
Institutional and Repository Deposit Policy
AOJSE fully supports self-archiving and green open access. Authors are encouraged to deposit the final published version of their article, also known as the version of record, in institutional repositories, subject repositories, and open-access databases immediately upon publication. There is no embargo period. Authors should always link back to the original article on the AOJSE website and include the full citation and DOI when depositing their work in any repository.
Digital Preservation and Archiving
AOJSE is committed to the long-term digital preservation of all published content to ensure its permanent availability and accessibility. The journal utilizes the Open Journal Systems (OJS) platform, which supports internationally recognized digital archiving standards. AOJSE encourages participation in archiving networks such as LOCKSS (Lots of Copies Keep Stuff Safe) and CLOCKSS (Controlled LOCKSS) to safeguard published articles against data loss and ensure continued access for future generations of researchers and readers.
Disclaimer
The views, opinions, findings, and conclusions expressed in articles published in AOJSE are solely those of the authors and do not necessarily reflect the official position, policy, or views of the journal, its editorial board, or its affiliated organizations. AOJSE makes every effort to ensure the accuracy and integrity of published content but does not accept legal responsibility for any errors, omissions, or claims arising from the use of information published in the journal.
Contact for Licensing Inquiries
For any questions regarding the licensing terms, copyright, or permitted use of content published in AOJSE, please contact the editorial office through the journal website at https://aojse.org. The editorial team will be happy to assist authors, readers, and institutions with any licensing-related queries.