Generative AI Architectures for Real-Time Transaction Processing

Authors

  • Jens Lehmann Author

Keywords:

Generative Artificial Intelligence, GenAI Enabled DevOps, Real Time Transaction Processing, Financial Services Pipelines, Black Swan Event Handling, Enterprise Architecture Theory, Cloud Native Architectures, Transaction Pipeline Scalability, DevOps Automation Practices, Real Time Data Processing, Operational Resilience Engineering, Non Functional Requirements Analysis, Intelligent Pipeline Monitoring, Adaptive System Design, Financial Infrastructure Modernization, Event Driven Architectures, High Volume Transaction Systems, AI Assisted Software Operations, Architectural Patterns And Trade Offs, Enterprise Grade Transaction Systems.

Abstract

Generative AI (GenAI) is opening new capabilities in software development and operations, reinforcing the broader shift toward DevOps principles and practices. At the same time, organizations across industries are under growing pressure to process data in real time. In financial services, this pressure is most acute in transaction pipelines, which must absorb volume spikes during Black Swan events despite comparatively modest average loads. GenAI-enabled DevOps offers promising ways to accelerate the design, implementation, and monitoring of such pipelines, yet its adoption in this context surfaces unresolved tensions. This paper reviews the existing literature, identifies gaps, draws on Generalized Enterprise Architecture Theory, and formulates research questions whose resolution can produce solutions that reconcile GenAI-enabled DevOps principles with the demands of real-time transaction processing. Toward that end, the paper lays the groundwork by specifying the functional and nonfunctional requirements of such pipelines, surveying GenAI-driven DevOps concepts, architectures, and operational paradigms relevant to financial systems, and outlining suitable design patterns along with their associated challenges and trade-offs.


References

1. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., ... Amodei, D. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.

2. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-T., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.

3. Mashetty, S., Malempati, M., Paleti, S., Adusupalli, B., & Singireddy, J. (2025). A Multidisciplinary Framework for AI and Data-Driven Transformation in Taxation, Insurance, Mortgage Financing, and Financial Advisory: Integrating Cloud Computing, Deep Learning, and Agentic AI for Community-Centric Economic Development. Insurance, Mortgage Financing, and Financial Advisory: Integrating Cloud Computing, Deep Learning, and Agentic AI for Community-Centric Economic Development.

4. Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., & Liu, P. J. (2020). Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140), 1–67.

5. Rae, J. W., Borgeaud, S., Cai, T., Millican, K., Hoffmann, J., Song, F., Aslanides, J., Henderson, S., Ring, R., Young, S., Rutherford, E., Hennigan, T., Menick, J., Cassirer, A., Powell, R., van den Driessche, G., Hendricks, L. A., Rauh, M., Huang, P.-S., ... Hassabis, D. (2021). Scaling language models: Methods, analysis & insights from training Gopher. arXiv preprint arXiv:2112.11446.

6. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). LoRA: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685.

7. Kummari, D. N., Singireddy, J., Sheelam, G. K., Nandan, B. P., Pandiri, L., Lakkarasu, P., & Dwaraka. (2025, August). Generative AI Models for Process Optimization in Semiconductor Wafer Design and Yield Prediction. In International Conference on Artificial Intelligence: Theory and Applications (pp. 220-233). Cham: Springer Nature Switzerland.

8. Fedus, W., Zoph, B., & Shazeer, N. (2022). Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120), 1–39.

9. Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvetkov, Y., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., ... Fiedel, N. (2022). PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240), 1–113.

10. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837.

11. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.

12. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2022). ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations.

13. Mashetty, S. (2025). LEVERAGING DEEP LEARNING, NEURAL NETWORKS, AND DATA ENGINEERING FOR INTELLIGENT MORTGAGE LOAN VALIDATION. INTERNATIONAL JOURNAL OF SOCIAL SCIENCE & INTERDISCIPLINARY RESEARCH ISSN: 2277-3630 Impact factor: 8.036, 14(04), 51-65.

14. Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Ré, C. (2022). FlashAttention: Fast and memory-efficient exact attention with IO-awareness. Advances in Neural Information Processing Systems, 35, 16344–16359.

15. Aminabadi, R. Y., Rajbhandari, S., Zhang, M., Awan, A. A., Li, C., Li, D., Zheng, E., Rasley, J., Smith, S., Ruwase, O., & He, Y. (2022). DeepSpeed inference: Enabling efficient inference of transformer models at unprecedented scale. Microsoft Research Technical Report MSR-TR-2022-21.

16. Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36, 68539–68551.

17. Touvron, H., Martin, L., Stone, K. R., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Brown, A., Bubeck, S., Cao, M., Caswell, I., Cecchi, M., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., ... Scialom, T. (2023). Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288.

18. Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., & Lample, G. (2023). LLaMA: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971.

19. OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., ... Zoph, B. (2023). GPT-4 technical report. arXiv preprint arXiv:2303.08774.

20. Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., & Stoica, I. (2023). Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles, 611–626.

21. Dao, T. (2023). FlashAttention-2: Faster attention with better parallelism and work partitioning. arXiv preprint arXiv:2307.08691.

22. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. Advances in Neural Information Processing Systems, 36.

23. Adusupalli, B., Malempati, M., Paleti, S., Mashetty, S., & Singireddy, J. (2025). Integrated financial ecosystems: AI-driven innovations in taxation, insurance, mortgage analytics, and community investment through cloud, big data, and advanced data engineering. Journal of Information Systems Engineering and Management, 10, 1103-1117.

24. Zhang, S., Zeng, X., Wu, Y., & Yang, Z. (2023). Harnessing scalable transactional stream processing for managing large language models. arXiv preprint arXiv:2307.08225.

25. Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., Casas, D. L., Bressand, F., Lengyel, G., Bour, G., Lample, G., ... Scao, T. L. (2024). Mixtral of experts. arXiv preprint arXiv:2401.04088.

26. Gemini Team. (2024). Gemini: A family of highly capable multimodal models. arXiv preprint arXiv:2312.11805.

27. Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Vaughan, A., ... Zhang, J. (2024). The Llama 3 herd of models. arXiv preprint arXiv:2407.21783.

28. Hammami, H., Baligand, L., & Petrovski, B. (2024). Fighting crime with Transformers: Empirical analysis of address parsing methods in payment data. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 201–212.

29. Bernardi, M. L., Casciani, A., Cimitile, M., & Marrella, A. (2024). Conversing with business process-aware large language models: The BPLLM framework. Journal of Intelligent Information Systems, 62, 1607–1629.

30. Karst, F. S., Chong, S.-Y., Antenor, A. A., Lin, E., Li, M. M., & Leimeister, J. M. (2024). Generative AI for banks: Benchmarks and algorithms for synthetic financial transaction data. arXiv preprint arXiv:2412.14730.

31. Recharla, M. (2024). Antioxidants, Biological Markers, Catalase, Glutathione Peroxidase, Chronic Periodontitis, Saliva, Smokeless tobacco, Smoker. Frontiers in Health Informatics, 13(8), 4999.

32. Murtaza, S. S., Nie, Y., Avan, E., Soni, U., Liao, W., Carnegie, A., Mathias, C. J., Jiang, J., & Wen, E. (2025). Implementing retrieval augmented generation technique on unstructured and structured data sources in a call center of a large financial institution. In Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 598–606.

33. Bhupathi, S. (2025). Role of databases in GenAI applications. arXiv preprint arXiv:2503.04847.

34. Ouafiq, E. M., & Saadane, R. (2025). Retrieval-augmented OLAP: Generative AI architecture for smart systems & equipment. Proceedings of the AAAI Symposium Series, 6(1), 304–312.

35. Ghali, M.-K., Farrag, A., Won, D., & Jin, Y. (2025). Enhancing knowledge retrieval with in-context learning and semantic search through generative AI. Knowledge-Based Systems, 311, 113047.

Additional Files

Published

2026-08-13

How to Cite

Generative AI Architectures for Real-Time Transaction Processing. (2026). The American Online Journal of Science and Engineering (AOJSE), 4(03). https://aojse.org/index.php/aojse/article/view/63

Most read articles by the same author(s)

Similar Articles

1-10 of 56

You may also start an advanced similarity search for this article.