CICD-Driven Real-Time Analytics on Azure Databricks

Authors

  • Ethan Williams Author

Keywords:

CI/CD for Data Engineering, DataOps Frameworks, MLOps Pipelines, DevSecOps Integration, Infrastructure as Code (IaC), Data Lakehouse Architecture, Data Mesh Adoption, Azure Databricks Pipelines, ETL Pipeline Automation, Data Pipeline Orchestration, Near Real-Time Analytics, Streaming Data Processing, Data Availability Optimization, Data Quality Governance, Secure Data Engineering, Observability in Data Systems, Cloud Data Platforms, Automated Data Deployment, Enterprise Data Engineering, Scalable Analytics Systems.

Abstract

An end-to-end Continuous Integration/Continuous Deployment (CI/CD) framework for data engineering is proposed, enabling automation from source control to development, testing, deployment, and monitoring. Supporting advanced development patterns like Infrastructure as Code (IaC), DataOps, and MLOps, these best practices enhance collaboration, streamline provisioning, and improve observability while ensuring consistent quality, security, and compliance. As the underpinning for a Zoomcar production-grade data engineering pipeline on Azure Databricks, the architecture supports a demand-based data engineering approach that favors data availability over access latency for Data Mesh adoption.

Cloud-based Data and Analytics platforms support advanced analytics, Artificial Intelligence, and Machine Learning on Near Real-Time Streaming data, enabling data-driven enterprise decision-making. Continuous Integration/Continuous Deployment (CI/CD) with DevSecOps is ubiquitous in Software Development and Application Development. Expanding CI/CD to ETL pipelines supporting MLOps and DataOps brings powerful, tested, and fast-productized code into Production for Cloud-Based Data Engineering Process. Data Availability on Demand supersedes Data Availability in Real-Time. A well-governed Data Lakehouse serving a Data Mesh architecture with a focus on Data Availability, Quality, and Security over Streaming Latency meets the demands of a Data Vision for a Data-Driven Enterprise.

References

1. Armbrust, M., Huai, Y., Liang, C., Xin, R., Zaharia, M., Franklin, M. J., & Ghodsi, A. (2021). Delta Lake: High-performance ACID table storage over cloud object stores. Proceedings of the VLDB Endowment, 13(12), 3411–3424.

2. Akidau, T., Chernyak, S., & Lax, R. (2021). Streaming Systems: The What, Where, When, and How of Large-Scale Data Processing (2nd ed.). O'Reilly Media.

3. Kolla, T., & Kolla, S. K. (2023). FHIR-Based Real-Time Healthcare Analytics using Unsupervised Learning. International Journal of Future Innovative Science and Technology (IJFIST), 6(6), 11751.

4. Chambers, B., & Zaharia, M. (2021). Spark: The Definitive Guide (Updated ed.). O'Reilly Media.

5. Raj, P., & Jaiswal, V. (2021). Azure Databricks Cookbook: Accelerate and Scale Real-Time Analytics Solutions Using the Apache Spark-Based Analytics Service. Packt Publishing.

6. Kleppmann, M. (2021). Designing Data-Intensive Applications (Updated ed.). O'Reilly Media.

7. Aitha, A. R. (2023). Cloud-Native Big Data AI/ML Framework for Risk Intelligence and Fraud Control in Banking and Insurance Ecosystems. Available at SSRN 6157967.

8. Karau, H., Warren, R., & Konwinski, A. (2021). High Performance Spark: Best Practices for Scaling and Optimizing Apache Spark (2nd ed.). O'Reilly Media.

9. Burns, B., Beda, J., & Hightower, K. (2022). Kubernetes: Up and Running (3rd ed.). O'Reilly Media.

10. Humble, J., & Farley, D. (2021). Continuous Delivery: Reliable Software Releases Through Build, Test, and Deployment Automation (Updated ed.). Addison-Wesley Professional.

11. Kim, G., Humble, J., Debois, P., & Willis, J. (2021). The DevOps Handbook (2nd ed.). IT Revolution Press.

12. Pereira, K., Vinagre, J., Alonso, A. N., Coelho, F., & Carvalho, M. (2022, September). Privacy-preserving machine learning in life insurance risk prediction. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases (pp. 44-52). Cham: Springer Nature Switzerland.

13. Turnbull, J. (2021). The Docker Book: Containerization Is the New Virtualization. James Turnbull.

14. Newman, S. (2021). Building Microservices (2nd ed.). O'Reilly Media.

15. Richards, M., & Ford, N. (2020). Fundamentals of Software Architecture: An Engineering Approach. O'Reilly Media.

16. Kleppmann, M., & Kreps, J. (2020). Kafka and stream processing for scalable real-time data platforms. IEEE Software, 37(5), 72–79.

17. Mattaparthi, R. (2023). Connected Fleet Intelligence: Edge-Centric Analytics and Computer Vision for Predictive Manufacturing and Asset Resilience. International Journal of Advanced Research in Computer Science & Technology (IJARCST), 6(5), 9077-9088.

18. Kreps, J., Narkhede, N., & Rao, J. (2020). Kafka: A distributed messaging system for modern real-time data pipelines. Communications of the ACM, 63(9), 74–83.

19. Zaharia, M., Xin, R., Wendell, P., Das, T., Armbrust, M., Dave, A., Meng, X., Rosen, J., Venkataraman, S., Franklin, M. J., Ghodsi, A., Gonzalez, J., Shenker, S., & Stoica, I. (2020). Apache Spark: A unified engine for big data processing. Communications of the ACM, 63(11), 56–65.

20. Marz, N., & Warren, J. (2021). Big Data: Principles and Best Practices of Scalable Realtime Data Systems (Updated ed.). Manning Publications.

21. Carbone, P., Katsifodimos, A., Ewen, S., Markl, V., Haridi, S., & Tzoumas, K. (2020). Apache Flink: Stream and batch processing in a single engine. IEEE Data Engineering Bulletin, 43(2), 28–38.

22. Grolinger, K., Higashino, W. A., Tiwari, A., & Capretz, M. A. M. (2020). Data management in cloud environments: NoSQL and NewSQL data stores. Journal of Cloud Computing, 9(1), 1–23.

23. Kolla, S. K. (2023). Learning Health Systems Machine Intelligence for Clinical Prediction and Healthcare Optimization. International Journal of Advanced Research in Computer Science & Technology (IJARCST), 6(2), 7955-7966.

24. Isah, H., Abughofa, T., Mahfuz, S., Ajerla, D., Zulkernine, F., & Khan, S. (2022). A survey of distributed data stream processing frameworks. IEEE Access, 10, 12345–12370.

25. Villamizar, M., Garcés, O., Castro, H., Verano, M., Salamanca, L., Casallas, R., & Gil, S. (2020). Infrastructure as Code for cloud-native software development and deployment: A systematic mapping study. Journal of Systems and Software, 167, 110599.

26. Sadalage, P. J., & Fowler, M. (2021). NoSQL Distilled: A Brief Guide to the Emerging World of Polyglot Persistence (2nd ed.). Addison-Wesley Professional.

27. Vavilapalli, V. K., Murthy, A. C., Douglas, C., Agarwal, S., Konar, M., Evans, R., Graves, T., Lowe, J., Shah, H., Seth, S., Saha, B., Curino, C., O'Malley, O., Radia, S., Reed, B., & Baldeschwieler, E. (2020). Apache Hadoop YARN: Yet another resource negotiator. Communications of the ACM, 63(1), 50–57.

28. Yandamuri, U. S. (2023). An Intelligent Analytics Framework Combining Big Data and Machine Learning for Business Forecasting. International Journal Of Finance, 36(6), 682-706.

29. Hellerstein, J. M., Ré, C., Schoppmann, F., Wang, D., Fratkin, E., Gorajek, A., Ng, K. S., Welton, C., Feng, X., Li, K., & Kumar, A. (2020). The MADlib analytics library: Machine learning in SQL. Proceedings of the VLDB Endowment, 13(12), 2837–2840.

30. Li, Y., Guo, L., Shen, H., & Li, X. (2020). Cloud-native big data analytics: Architectures, techniques, and applications. Future Generation Computer Systems, 107, 1004–1017.

31. Ebert, C., Gallardo, G., Hernantes, J., & Serrano, N. (2021). DevOps. IEEE Software, 38(2), 94–100.

32. Davuluri, P. N. AI-Augmented Sanctions Screening: Enhancing Accuracy and Latency in Real Time Compliance Systems.

33. Rahman, A., Williams, L., & Ståhl, D. (2021). Continuous integration practices and software quality: A systematic literature review. Information and Software Technology, 137, 106588.

34. Shahin, M., Ali Babar, M., & Zhu, L. (2020). Continuous integration, delivery, and deployment: A systematic review on approaches, tools, challenges, and practices. IEEE Access, 8, 170255–170291.

35. Hassan, S., Bahsoon, R., & Kazman, R. (2021). Microservices and cloud-native architectures: Trends and research directions. IEEE Software, 38(5), 24–31.

36. Nagabhyru, K. C., & Engineer, S. D. (2023). Unifying Data Engineering and Machine Learning Pipelines: An Enterprise Roadmap to Automated Model Deployment.

37. Casado, M., Foster, N., & Guha, A. (2020). Software-defined networking and cloud computing: Current trends and future directions. Communications of the ACM, 63(7), 64–73.

38. Stonebraker, M., Abadi, D. J., DeWitt, D. J., Madden, S., Paulson, E., Pavlo, A., & Rasin, A. (2020). MapReduce and parallel database systems: Friends or foes? Communications of the ACM, 63(8), 64–71.

39. Ghosh, S., & Grolinger, K. (2022). DataOps: Continuous integration and delivery for data analytics pipelines. IEEE Access, 10, 67542–67561.

40. Inala, R. Designing Scalable Technology Architectures for Customer Data in Group Insurance and Investment Platforms.

41. Bogner, J., Fritzsch, J., Wagner, S., & Zimmermann, A. (2021). Microservices in industry: Insights into technologies, characteristics, and software quality. IEEE International Conference on Software Architecture Companion, 187–195.

42. Grus, J. (2021). Data Science from Scratch: First Principles with Python (2nd ed.). O'Reilly Media.

43. Beyer, B., Jones, C., Petoff, J., & Murphy, N. R. (2020). Site Reliability Engineering: How Google Runs Production Systems (Updated ed.). O'Reilly Media.

44. Bandi, V. D. V. K. (2023). MLOps frameworks for reliable model deployment in cloud data platforms. Journal of Artificial Intelligence and Big Data, 3(1), 84-101.

45. Forsgren, N., Humble, J., & Kim, G. (2021). Accelerate: The Science of Lean Software and DevOps (Updated ed.). IT Revolution Press.

46. Narkhede, N., Shapira, G., & Palino, T. (2021). Kafka: The Definitive Guide (2nd ed.). O'Reilly Media.

47. Kleppmann, M. (2022). Data-intensive distributed systems: Principles and modern applications. ACM Computing Surveys, 55(8), 1–36.

48. Reddy, V. A. R. (2022). Data-Driven Healthcare Operations: Architecting Unified Member, Provider, and Claims Intelligence Platforms. International Journal of Science, Research and Technology, 5(5), 8511-8521.

49. Armbrust, M., Das, T., Davidson, A., Ghodsi, A., Or, A., Rosen, J., Xin, R., Zaharia, M., & Franklin, M. J. (2022). Lakehouse architecture: Combining data lakes and data warehouses. Communications of the ACM, 65(7), 72–81.

50. Verma, A., Pedrosa, L., Korupolu, M., Oppenheimer, D., Tune, E., & Wilkes, J. (2020). Large-scale cluster management at Google with Borg. Communications of the ACM, 63(2), 66–74.

51. Eismann, S., Grohmann, J., Herbst, N. R., Kounev, S., & Abad, C. L. (2021). A review of serverless computing for cloud applications. IEEE Transactions on Cloud Computing, 9(3), 1202–1219.

52. Mangala, N. (2022). Implementing Databricks Unity Catalog For Centralized Data Governance In Multi-Business-Unitenterprises. Journal of International Crisis and Risk Communication Research, 101-122.

53. Armbrust, M., Ghodsi, A., Xin, R., Zaharia, M., Franklin, M. J., & Dave, A. (2023). The lakehouse architecture for modern data platforms. Proceedings of the VLDB Endowment, 16(5), 1203–1216.

54. Zaharia, M., Chen, A., Davidson, A., Ghodsi, A., Hong, S. A., Konwinski, A., Murching, S., Nykodym, T., Ogilvie, P., Parkhe, M., & Xin, R. (2023). Lakehouse: A new generation of open platforms that unify data warehousing and advanced analytics. Communications of the ACM, 66(8), 62–71.

55. Gottimukkala, V. R. R. (2020). Energy-Efficient Design Patterns for Large-Scale Banking Applications Deployed on AWS Cloud. power, 9(12).

56. Kreps, J. (2022). Data streaming and event-driven architectures for modern cloud applications. IEEE Software, 39(5), 34–41.

57. Grolinger, K., Higashino, W. A., Tiwari, A., & Capretz, M. A. M. (2022). Big data analytics in cloud computing: A survey of techniques, applications, and tools. Journal of Cloud Computing, 11(1), 1–28.

58. Ebert, C., & Gallardo, G. (2022). DevOps for digital transformation and cloud-native software engineering. IEEE Software, 39(4), 16–23.

59. Forsgren, N., Humble, J., & Kim, G. (2022). Continuous delivery performance in cloud software development. IEEE Software, 39(6), 78–85.

60. Inala, R. Advancing Group Insurance Solutions Through Ai-Enhanced Technology Architectures And Big Data Insights.

61. Rahman, A., Williams, L., & Ståhl, D. (2022). Improving software quality through continuous integration and deployment practices. Information and Software Technology, 145, 106822.

62. Shahin, M., Ali Babar, M., & Zhu, L. (2023). Advances in continuous integration, delivery, and deployment: Trends and future research. Journal of Systems and Software, 197, 111560.

63. Narkhede, N., Shapira, G., & Palino, T. (2022). Event streaming architectures using Apache Kafka for enterprise analytics. IEEE Internet Computing, 26(5), 30–39.

64. Li, Y., Guo, L., Shen, H., & Li, X. (2023). Cloud-native data analytics: Challenges and opportunities for enterprise applications. Future Generation Computer Systems, 140, 255–268.

65. Mangalampalli, B. M. (2022). Automated Invoice Validation Systems Using Advanced SQL Analytics in Healthcare Insurance. Front Health Inform, 11.

66. Bogner, J., Fritzsch, J., Wagner, S., & Zimmermann, A. (2022). Cloud-native microservices: Current practices and future directions. IEEE Software, 39(2), 88–96.

67. Eismann, S., Grohmann, J., Herbst, N. R., Kounev, S., & Abad, C. L. (2022). Serverless computing for scalable cloud analytics: A systematic review. IEEE Transactions on Cloud Computing, 10(4), 2384–2398.

68. Burns, B., Beda, J., & Hightower, K. (2023). Kubernetes orchestration for cloud-native data platforms. IEEE Cloud Computing, 10(2), 24–33.

69. Carbone, P., Katsifodimos, A., Markl, V., Haridi, S., & Tzoumas, K. (2022). Stream processing with Apache Flink: Recent developments and applications. IEEE Data Engineering Bulletin, 45(3), 30–44.

70. Isah, H., Abughofa, T., Mahfuz, S., Ajerla, D., Zulkernine, F., & Khan, S. (2023). Distributed stream processing frameworks for real-time big data analytics: An updated survey. IEEE Access, 11, 47251–47279.

71. Ghosh, S., & Grolinger, K. (2023). DataOps for enterprise-scale machine learning and analytics pipelines. IEEE Access, 11, 87354–87372.

72. Davuluri, P. N. Integrating Artificial Intelligence into Event-Driven Financial Crime Compliance Platforms.

73. Verma, A., Pedrosa, L., Korupolu, M., Oppenheimer, D., Tune, E., & Wilkes, J. (2022). Resource management strategies for large-scale cloud clusters. Communications of the ACM, 65(4), 74–83.

74. Hassan, S., Bahsoon, R., & Kazman, R. (2023). Engineering cloud-native software systems with microservices and DevOps. IEEE Software, 40(3), 58–66.

75. Kleppmann, M. (2023). Data-intensive applications in the era of cloud-native architectures. ACM Computing Surveys, 56(2), 1–35.

76. Villamizar, M., Garcés, O., Castro, H., Salamanca, L., Casallas, R., & Gil, S. (2023). Infrastructure as code for continuous deployment of cloud-native applications: Recent advances and industrial practices. Journal of Systems and Software, 197, 111569.

Additional Files

Published

2023-12-14

How to Cite

CICD-Driven Real-Time Analytics on Azure Databricks. (2023). The American Online Journal of Science and Engineering (AOJSE), 1(01). https://aojse.org/index.php/aojse/article/view/26

Similar Articles

1-10 of 23

You may also start an advanced similarity search for this article.