Hybrid AI for Automated Clinical Risk Coding

Authors

  • Dileep Valiki Author

Keywords:

Value-Based Healthcare, Healthcare Risk Adjustment, Hierarchical Condition Categories (HCC), International Classification of Diseases (ICD), Medicare Advantage Risk Models, Generative Artificial Intelligence in Healthcare, Large Language Models for Clinical NLP, Automated HCC Code Identification, Clinical Note Information Extraction, Hybrid LLM–NLP Architectures, Embedding-Based Clinical Classification, Unstructured Clinical Data Processing, Healthcare Payment Optimization, Domain-Specific Language Models, AI-Driven Medical Coding.

Abstract

The long-overdue transition of the United States health system toward value-based care has triggered a new demand for healthcare risk adjustment. The chronic medical complexities of at-risk populations are captured and represented through clinical concepts from the International Classification of Diseases (ICD). These codes are prevalent in payment models such as the Medicare Advantage program and the Affordable Care Act marketplace. Medicare Advantage payers receive funding from the Medicare program based on a patient-level risk adjustment prediction model, which is based on Hierarchical Condition Categories (HCC). The accurate and efficient identification of HCC codes in unstructured clinical notes could reduce operational costs while directly benefiting the medical community and patients. However, few solutions address this problem. Generative AI can help address these emerging needs by automating the complex natural language processing (NLP) tasks involved in HCC code identification.

This research study proposes a novel coding architecture that employs generative AI for automatic HCC code identification from clinical notes. A hybrid large language model-natural language processing (LLM-NLP) generative AI architecture is designed; it includes a large language model that generates a set of candidate HCC codes, which is filtered based on the restricted natural language processing of classifiers using a new embedding approach. The novel architecture is applicable to domain-specific tasks requiring extracting relevant concepts from a very large number of entity classes.

References

[1] Soroush, A. (2024). Large language models are poor medical coders: Benchmarking of medical code querying. NEJM AI, 1(2).

[2] Yoo, Y., et al. (2025). How to leverage large language models for automatic ICD coding: Fine-tuning guidelines and a robust framework. Computers in Biology and Medicine, 176, 108322.

[3] Puts, S., et al. (2025). Developing an ICD-10 coding assistant: Pilot study using retrieval-augmented generation–based methods. Journal of Medical Internet Research, 27, e11835781.

[4] Wu, J., et al. (2024). Transparency in language models for automated medical coding. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics.

[5] Zhuang, Y., et al. (2025). Evaluation of an automated analysis system using medical text: Prompt learning framework for ICD coding. JMIR Medical Informatics, 13, e63020.

[6] Bhutto, S. R., et al. (2024). Automatic ICD-10-CM coding via lambda-scaled attention for long clinical notes. Artificial Intelligence in Medicine, 151, 102700.

[7] Li, F., & Yu, H. (2020). ICD coding from clinical text using multi-filter residual convolutional neural network. Artificial Intelligence in Medicine, 103, 101748.

[8] Vu, T., Nguyen, D. Q., & Nguyen, A. (2020). A label attention model for ICD coding from clinical text. In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence (IJCAI).

[9] Mullenbach, J., Wiegreffe, S., Duke, J., Sun, J., & Eisenstein, J. (2018). Explainable prediction of medical codes from clinical text. In Proceedings of NAACL-HLT 2018. Association for Computational Linguistics.

[10] Liu, L., et al. (2022). Hierarchical label-wise attention transformer model for explainable prediction of ICD codes from clinical documents. Journal of Biomedical Informatics, 128, 104036.

[11] Niu, K., et al. (2023). Retrieve and rerank for automated ICD coding via contrastive learning. Journal of Biomedical Informatics, 146, 104317.

[12] Zhang, Y., et al. (2025). Automated ICD coding using deep learning and language models: A systematic review. Journal of Biomedical Informatics, 156, 104650.

[13] Sheikhalishahi, S., Miotto, R., Dudley, J. T., Lavelli, A., Rinaldi, F., & Osmani, V. (2019). Natural language processing of clinical notes on chronic diseases: Systematic review. JMIR Medical Informatics, 7(2), e12239.

[14] Wu, S., Roberts, K., Datta, S., et al. (2020). Deep learning in clinical natural language processing: A methodical review. Journal of the American Medical Informatics Association, 27(3), 457–470.

[15] Wehbe, R. M., et al. (2021). Deep learning in clinical natural language processing: A systematic review. Journal of the American Medical Informatics Association, 28(2), 1–15.

[16] Johnson, A. E. W., Pollard, T. J., Shen, L., et al. (2016). MIMIC-III, a freely accessible critical care database. Scientific Data, 3, 160035.

[17] Johnson, A. E. W., Stone, D. J., Celi, L. A., & Pollard, T. J. (2021). The MIMIC-IV clinical database. Scientific Data, 8, 257.

[18] Yang, X., et al. (2022). A large language model for electronic health records. npj Digital Medicine, 5, 194.

[19] Peng, C., et al. (2023). A study of generative large language model for medical research and healthcare. npj Digital Medicine, 6, 210.

[20] Luo, R., Sun, L., Xia, Y., et al. (2022). BioGPT: Generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 23(6), bbac409.

[21] Lee, J., Yoon, W., Kim, S., et al. (2020). BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4), 1234–1240.

[22] Gu, Y., Tinn, R., Cheng, H., et al. (2021). Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare, 3(1), 1–23.

[23] Beltagy, I., Lo, K., & Cohan, A. (2019). SciBERT: A pretrained language model for scientific text. In Proceedings of EMNLP-IJCNLP 2019. Association for Computational Linguistics.

[24] Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT 2019. Association for Computational Linguistics.

[25] Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.

[26] Wolf, T., Debut, L., Sanh, V., et al. (2020). Transformers: State-of-the-art natural language processing. In Proceedings of EMNLP 2020: System Demonstrations. Association for Computational Linguistics.

[27] Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.

[28] Ji, Z., Lee, N., Frieske, R., et al. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1–38.

[29] Kung, T. H., Cheatham, M., Medenilla, A., et al. (2023). Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digital Health, 2(2), e0000198.

[30] Gilson, A., Safranek, C. W., Huang, T., et al. (2023). How does ChatGPT perform on the United States Medical Licensing Examination? JMIR Medical Education, 9, e45312.

[31] Singhal, K., et al. (2023). Large language models encode clinical knowledge. Nature, 620, 172–180.

[32] Rajkomar, A., Oren, E., Chen, K., et al. (2018). Scalable and accurate deep learning with electronic health records. npj Digital Medicine, 1, 18.

[33] Shickel, B., Tighe, P. J., Bihorac, A., & Rashidi, P. (2018). Deep EHR: A survey of recent advances in deep learning techniques for electronic health record analysis. IEEE Journal of Biomedical and Health Informatics, 22(5), 1589–1604.

[34] Dernoncourt, F., Lee, J. Y., Uzuner, O., & Szolovits, P. (2017). De-identification of patient notes with recurrent neural networks. Journal of the American Medical Informatics Association, 24(3), 596–606.

[35] Lehman, E., et al. (2021). Does BERT pretrained on clinical text leak private information? In Proceedings of ACL-IJCNLP 2021. Association for Computational Linguistics.

[36] El Emam, K., & Dankar, F. K. (2008). Protecting privacy using k-anonymity. Journal of the American Medical Informatics Association, 15(5), 627–637.

[37] Office of the National Coordinator for Health Information Technology. (2020). United States Core Data for Interoperability (USCDI) v1. U.S. Department of Health and Human Services.

[38] Sittig, D. F., & Singh, H. (2016). A socio-technical approach to preventing, mitigating, and recovering from EHR-related safety hazards. Journal of the American Medical Informatics Association, 23(4), 641–647.

[39] Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). Why should I trust you? Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM.

[40] Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. Advances in Neural Information Processing Systems, 30, 4765–4774.

[41] Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1, 206–215.

[42] Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2021). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54(6), 1–35.

[43] Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453.

[44] Rajkomar, A., Hardt, M., Howell, M. D., Corrado, G., & Chin, M. H. (2018). Ensuring fairness in machine learning to advance health equity. Annals of Internal Medicine, 169(12), 866–872.

[45] Chen, I. Y., Joshi, S., Ghassemi, M., & Ranganath, R. (2021). Probabilistic machine learning for healthcare. Annual Review of Biomedical Data Science, 4, 393–419.

[46] Pope, G. C., Kautter, J., Ellis, R. P., et al. (2004). Risk adjustment of Medicare capitation payments using the CMS-HCC model. Health Care Financing Review, 25(4), 119–141.

[47] Ellis, R. P., & McGuire, T. G. (2007). Predictability and predictiveness in health care spending. Journal of Health Economics, 26(1), 25–48.

[48] Newhouse, J. P., Buntin, M. B., & Chapman, J. D. (1997). Risk adjustment and Medicare: Taking a closer look. Health Affairs, 16(5), 26–43.

[49] van de Ven, W. P. M. M., & Ellis, R. P. (2000). Risk adjustment in competitive health plan markets. In A. J. Culyer & J. P. Newhouse (Eds.), Handbook of Health Economics (Vol. 1, pp. 755–845). Elsevier.

[50] McGuire, T. G., Glazer, J., Newhouse, J. P., et al. (2013). Paying for performance in health insurance: Selection, incentives, and risk adjustment. Journal of Health Economics, 32(5), 901–912.

[51] Rose, S. (2016). A machine learning framework for plan payment risk adjustment. Health Services Research, 51(6), 2358–2374.

[52] Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning (2nd ed.). Springer.

[53] Lafferty, J., McCallum, A., & Pereira, F. (2001). Conditional random fields: Probabilistic models for segmenting and labeling sequence data. In Proceedings of the 18th International Conference on Machine Learning (ICML).

[54] Hochreiter, S., & Schmidhuber, J. (1997). Long short-term memory. Neural Computation, 9(8), 1735–1780.

[55] Cho, K., van Merriënboer, B., Gulcehre, C., et al. (2014). Learning phrase representations using RNN encoder-decoder for statistical machine translation. In Proceedings of EMNLP 2014. Association for Computational Linguistics.

[56] Pennington, J., Socher, R., & Manning, C. D. (2014). GloVe: Global vectors for word representation. In Proceedings of EMNLP 2014. Association for Computational Linguistics.

[57] Peters, M. E., Neumann, M., Iyyer, M., et al. (2018). Deep contextualized word representations. In Proceedings of NAACL-HLT 2018. Association for Computational Linguistics.

[58] Manning, C. D., Surdeanu, M., Bauer, J., Finkel, J., Bethard, S. J., & McClosky, D. (2014). The Stanford CoreNLP natural language processing toolkit. In Proceedings of ACL 2014: System Demonstrations. Association for Computational Linguistics.

[59] Uzuner, O., South, B. R., Shen, S., & DuVall, S. L. (2011). 2010 i2b2/VA challenge on concepts, assertions, and relations in clinical text. Journal of the American Medical Informatics Association, 18(5), 552–556.

[60] Henry, S., Buchan, K., Filannino, M., Stubbs, A., & Uzuner, O. (2020). 2018 n2c2 shared task on adverse drug events and medication extraction in electronic health records. Journal of the American Medical Informatics Association, 27(1), 3–12.

[61] Chapman, W. W., Nadkarni, P. M., Hirschman, L., D’Avolio, L. W., Savova, G. K., & Uzuner, O. (2011). Overcoming barriers to NLP for clinical text: The role of shared tasks and challenges. Journal of the American Medical Informatics Association, 18(5), 540–543.

[62] Savova, G. K., Masanz, J. J., Ogren, P. V., et al. (2010). Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): Architecture, component evaluation and applications. Journal of the American Medical Informatics Association, 17(5), 507–513.

[63] Bodenreider, O. (2004). The Unified Medical Language System (UMLS): Integrating biomedical terminology. Nucleic Acids Research, 32(Database issue), D267–D270.

[64] World Health Organization. (2019). International statistical classification of diseases and related health problems (11th ed.). World Health Organization.

[65] Steindel, S. J. (2010). International classification of diseases, 10th edition, clinical modification and procedure coding system: Descriptive overview. Clinical Chemistry, 56(6), 1053–1054.

[66] Berwick, D. M., Nolan, T. W., & Whittington, J. (2008). The triple aim: Care, health, and cost. Health Affairs, 27(3), 759–769.

[67] Porter, M. E. (2010). What is value in health care? New England Journal of Medicine, 363(26), 2477–2481.

[68] Mandel, J. C., Kreda, D. A., Mandl, K. D., Kohane, I. S., & Ramoni, R. B. (2016). SMART on FHIR: A standards-based, interoperable apps platform for electronic health records. Journal of the American Medical Informatics Association, 23(5), 899–908.

[69] Payne, T. H., Corley, S., Cullen, T. A., Gandhi, T. K., Harrington, L., Kuperman, G. J., et al. (2018). Report of the AMIA EHR 2020 task force on the status and future direction of EHRs. Journal of the American Medical Informatics Association, 22(5), 110–112.

[70] Raghupathi, W., & Raghupathi, V. (2014). Big data analytics in healthcare: Promise and potential. Health Information Science and Systems, 2, 3.

[71] Jensen, P. B., Jensen, L. J., & Brunak, S. (2012). Mining electronic health records: Towards better research applications and clinical care. Nature Reviews Genetics, 13(6), 395–405.

[72] Topol, E. (2019). Deep medicine: How artificial intelligence can make healthcare human again. Basic Books.

[73] Shortliffe, E. H., & Cimino, J. J. (2021). Biomedical informatics: Computer applications in health care and biomedicine (5th ed.). Springer.

[74] Fielding, R. T. (2000). Architectural styles and the design of network-based software architectures (Doctoral dissertation). University of California.

[75] Newman, S. (2015). Building microservices. O’Reilly Media.

[76] Pahl, C. (2015). Containerization and microservices for modern software engineering. IEEE Cloud Computing, 2(3), 24–31.

[77] Kreps, J., Narkhede, N., & Rao, J. (2011). Kafka: A distributed messaging system for log processing. In Proceedings of the NetDB Workshop.

[78] Akidau, T., et al. (2015). The dataflow model: A practical approach to balancing correctness, latency, and cost in massive-scale, unbounded, out-of-order data processing. Proceedings of the VLDB Endowment, 8(12), 1792–1803.

[79] Stonebraker, M., Çetintemel, U., & Zdonik, S. (2005). The 8 requirements of real-time stream processing. ACM SIGMOD Record, 34(4), 42–47.

[80] Mell, P., & Grance, T. (2011). The NIST definition of cloud computing. National Institute of Standards and Technology.

[81] Abouelmehdi, K., Beni-Hessane, A., & Khaloufi, H. (2018). Big healthcare data: Preserving security and privacy. Journal of Big Data, 5, 1–18.

[82] Wilkowska, W., & Ziefle, M. (2012). Privacy and data security in e-health: Requirements from the user’s perspective. Health Policy and Technology, 1(3), 119–131.

[83] Khezr, S., Moniruzzaman, M., Yassine, A., & Benlamri, R. (2019). Blockchain technology in healthcare: A comprehensive review and directions for future research. IEEE Access, 7, 173174–173197.

[84] Zhang, P., White, J., Schmidt, D. C., Lenz, G., & Rosenbloom, S. T. (2018). FHIRChain: Applying blockchain to securely and scalably share clinical data. Computational and Structural Biotechnology Journal, 16, 267–278.

[85] Meystre, S. M., Lovis, C., Bürkle, T., Tognola, G., Budrionis, A., & Lehmann, C. U. (2017). Clinical data reuse or secondary use: Current status and potential future progress. Yearbook of Medical Informatics, 26(1), 38–52.

[86] Greenhalgh, T., Wherton, J., Papoutsi, C., et al. (2017). Beyond adoption: A new framework for theorizing and evaluating nonadoption, abandonment, and challenges to the scale-up, spread, and sustainability of health and care technologies. Journal of Medical Internet Research, 19(11), e367.

[87] Bates, D. W., Cohen, M., Leape, L. L., Overhage, J. M., Shabot, M. M., & Sheridan, T. (2001). Reducing the frequency of errors in medicine using information technology. Journal of the American Medical Informatics Association, 8(4), 299–308.

[88] Mandl, K. D., & Kohane, I. S. (2015). No small change for the health information economy. New England Journal of Medicine, 372(14), 1278–1281.

[89] Sinsky, C., Colligan, L., Li, L., et al. (2016). Allocation of physician time in ambulatory practice: A time and motion study. Annals of Internal Medicine, 165(11), 753–760.

[90] Haux, R. (2010). Medical informatics: Past, present, future. International Journal of Medical Informatics, 79(9), 599–610.

Additional Files

Published

2025-06-26

How to Cite

Hybrid AI for Automated Clinical Risk Coding. (2025). The American Online Journal of Science and Engineering (AOJSE), 3(02). https://aojse.org/index.php/aojse/article/view/7

Similar Articles

1-10 of 23

You may also start an advanced similarity search for this article.