Neuro-Behavioral Fusion for Predictive Tax Fraud Detection in Cloud-Based Government Systems

Authors

  • Bhasker Katta Author

DOI:

https://doi.org/10.5281/zenodo.20742667

Keywords:

Cloud computing, behavioural analytics, detect, fraud, government systems, hybrid architectures, hybrid deep learning, predictive, tax, Twitter.

Abstract

Accurate prediction of tax fraud fosters favorable taxpayer-business relationships, streamlines tax authorities' operations, and optimally directs government investments to enhance public services. Reliable prediction of fraudulent taxpayers, however, requires careful selection of assessed features because of pronounced privacy, ethical, and security concerns associated with government-related data. Furthermore, limited historical records are available because fraud often remains undetected. Consequently, effective machine learning-based predictive models must deploy hybrid architectures capable of learning from different types of features, integrating feature engineering with representation learning, and extracting fraud behavioral patterns during training. Moreover, tax fraud detection is a pattern-discovery problem with critical imbalance between the positive and negative classes, necessitating the use of behaviorally informative indicators that optimize predictive performance, fairness, and robustness.

Predictive models should therefore integrate hybrid deep learning representations with behavioral-based tax fraud indicators for risk detection. Cloud-enabled government data ecosystems provide abundant data sources for detecting and predicting all forms of tax fraud, and during operation can be simulated to include as many behaviors that are indicators of tax fraud as possible. Predictive models can thus be trained and tested using a behavioral pattern-discovery approach, forming the foundation for hybrid deep learning-based models that use a dual-tower architecture to employ both behavioral indicators and independently engineered features to optimize detection risk.

References

[1] Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., Kudlur, M., Levenberg, J., Monga, R., Moore, S., Murray, D. G., Steiner, B., Tucker, P., Vasudevan, V., Warden, P., … Zheng, X. (2016). TensorFlow: A system for large-scale machine learning. In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’16) (pp. 265–283). USENIX Association.

[2] Akoglu, L., Tong, H., & Koutra, D. (2015). Graph based anomaly detection and description: A survey. Data Mining and Knowledge Discovery, 29(3), 626–688.

[3] Arrieta, A. B., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115.

[4] Audet, A., & Neyshabur, B. (2023). Implicit bias of optimization algorithms in deep learning: A survey. Journal of Machine Learning Research, 24, 1–89.

[5] Bahdanau, D., Cho, K., & Bengio, Y. (2015). Neural machine translation by jointly learning to align and translate. In Proceedings of the 3rd International Conference on Learning Representations (ICLR).

[6] Bao, W., Yue, J., & Rao, Y. (2017). A deep learning framework for financial time series using stacked autoencoders and long-short term memory. PLOS ONE, 12(7), e0180944.

[7] Barredo Arrieta, A., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable Artificial Intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115.

[8] Biecek, P. (2018). DALEX: Explainers for complex predictive models in R. Journal of Machine Learning Research, 19, 1–5.

[9] Bishop, C. M. (2006). Pattern recognition and machine learning. Springer.

[10] Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32.

[11] Budach, L., Feuerpfeil, J., Ihde, N., Nathansen, A., Noack, D., Patzlaff, N., Naumann, F., & Spiekermann, M. (2022). The effects of data augmentation on deep learning for tabular data. Proceedings of the VLDB Endowment, 15(8), 1580–1593.

[12] Carcillo, F., Le Borgne, Y. A., Caelen, O., Kessaci, Y., Oblé, F., & Bontempi, G. (2019). Combining unsupervised and supervised learning in credit card fraud detection. Information Sciences, 557, 317–331.

[13] Chandola, V., Banerjee, A., & Kumar, V. (2009). Anomaly detection: A survey. ACM Computing Surveys, 41(3), 1–58.

[14] Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM.

[15] Christodoulou, E., Ma, J., Collins, G. S., Steyerberg, E. W., Verbakel, J. Y., & Van Calster, B. (2019). A systematic review shows no performance benefit of machine learning over logistic regression for clinical prediction models. Journal of Clinical Epidemiology, 110, 12–22.

[16] Cloudera. (2020). Modern data architecture and governance in the cloud. Cloudera White Paper.

[17] Cohen, W. W., Ravikumar, P., & Fienberg, S. E. (2003). A comparison of string distance metrics for name-matching tasks. In Proceedings of the IJCAI Workshop on Information Integration on the Web (pp. 73–78).

[18] Cortes, C., & Vapnik, V. (1995). Support-vector networks. Machine Learning, 20(3), 273–297.

[19] Dehghani, M., Shakeri, H., & Zohrevand, A. (2021). Fraud detection in financial statements using deep learning and ensemble approaches: A systematic review. Expert Systems with Applications, 168, 114–168.

[20] Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT 2019 (pp. 4171–4186). Association for Computational Linguistics.

[21] Doersch, C. (2016). Tutorial on variational autoencoders. arXiv preprint arXiv:1606.05908.

[22] Dua, D., & Graff, C. (2019). UCI machine learning repository. University of California, Irvine.

[23] Esteva, A., Robicquet, A., Ramsundar, B., Kuleshov, V., DePristo, M., Chou, K., Cui, C., Corrado, G., Thrun, S., & Dean, J. (2019). A guide to deep learning in healthcare. Nature Medicine, 25(1), 24–29.

[24] Feurer, M., Klein, A., Eggensperger, K., Springenberg, J. T., Blum, M., & Hutter, F. (2015). Efficient and robust automated machine learning. In Advances in Neural Information Processing Systems (pp. 2962–2970).

[26] Friedler, S. A., Scheidegger, C., Venkatasubramanian, S., Choudhary, S., Hamilton, E. P., & Roth, D. (2019). A comparative study of fairness-enhancing interventions in machine learning. In Proceedings of FAT* 2019 (pp. 329–338). ACM.

[27] Gama, J., Žliobaitė, I., Bifet, A., Pechenizkiy, M., & Bouchachia, A. (2014). A survey on concept drift adaptation. ACM Computing Surveys, 46(4), 1–37.

[26] Garfinkel, S. L. (2015). De-identification of personal information. NIST Internal Report.

[27] Ghorbani, A., Wexler, J., Zou, J., & Kim, B. (2019). Towards automatic concept-based explanations. In Advances in Neural Information Processing Systems (pp. 9277–9286).

[28] Gilpin, L. H., Bau, D., Yuan, B. Z., Bajwa, A., Specter, M., & Kagal, L. (2018). Explaining explanations: An overview of interpretability of machine learning. In 2018 IEEE 5th International Conference on Data Science and Advanced Analytics (DSAA) (pp. 80–89). IEEE.

[29] Goodfellow, I., Bengio, Y., & Courville, A. (2016). Deep learning. MIT Press.

[30] Gopalan, P. K., Hofman, J. M., & Blei, D. M. (2012). Scalable recommendation with Poisson factorization. In Proceedings of the 28th Conference on Uncertainty in Artificial Intelligence (pp. 326–335).

[31] Guo, C., & Berkhahn, F. (2016). Entity embeddings of categorical variables. arXiv preprint arXiv:1604.06737.

[32] Hamilton, W. L. (2020). Graph representation learning. Morgan & Claypool.

[33] Han, J., Kamber, M., & Pei, J. (2011). Data mining: Concepts and techniques (3rd ed.). Morgan Kaufmann.

Additional Files

Published

2023-12-19

How to Cite

Neuro-Behavioral Fusion for Predictive Tax Fraud Detection in Cloud-Based Government Systems. (2023). The American Online Journal of Science and Engineering (AOJSE), 1(01). https://doi.org/10.5281/zenodo.20742667

Similar Articles

1-10 of 21

You may also start an advanced similarity search for this article.