Multi-Agent RL for Value-Based Population Health Optimization

Authors

  • Vinod Battapothu Author

DOI:

https://doi.org/10.5281/zenodo.20744351

Keywords:

Agentic Artificial Intelligence, Multi-Agent Reinforcement Learning (MARL), Population Health Management, Risk- and Performance-Based Contracts, Healthcare Payment Redistribution, Actor–Network Healthcare Models, Complex Adaptive Healthcare Systems, Agent-Based Decision Support, Health Information Systems Innovation, Risk-Adjusted Outcome Optimization, Healthcare Econometrics, Decentralized Healthcare Governance, AI-Driven

Abstract

Agentic artificial intelligence (AI) systems, including those realized through multi-agent reinforcement learning (MARL), are emerging in multiple domains, including information filtering, navigation, gaming, and robotics. These novel approaches may also be relevant to healthcare delivery and policy, as MARL is well suited to complex strategic problems that Richard Thaler characterized as “makeshift solutions to impossibly complex problems.” Population health management is such an arena, particularly in optimizing risk- and performance-based contracts, payment redistribution across agents and cadre types, and performance-related benchmarks for an actor-network consisting of several competing provider-organizations functioning in a shared ecosystem. Health information systems, including such agentic, MARL-based systems, could help realize “better, faster, cheaper, and ‘kinder’ healthcare” by improving risk-adjusted health and cost outcomes across multiple resident cohorts, thereby more effectively realizing the goals of modern econometrics in medicine and public health.

Current culture tends to represent healthcare decision-making as a Black Box or an informationally centralized “big brain” where supervision assures correctness and compliance with Thaler’s ideal of omniscience for coordinating activities. Yet actual healthcare organizations are not Black Boxes, much less a single Big Brain, corporate or otherwise. Decision makers are often overpromoted, misinformed, late in their decisions, veto the best ideas of the organization, or are unable to manage rich information effectively. Supporting such overloaded Decision Makers-MD, using training and expertise gained trying to wisely harvest large but incomplete information- is a hard problem. Because the decisions should be seen as Agents making active decisions, it is viewed in terms of Agent-based Systems, or the study of multi-oriented organizations as complex adaptive systems.

References

[1] Agarwal, A., Dudík, M., Langford, J., & Li, L. (2014). A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web (pp. 661–670). ACM.

[2] Athey, S., & Imbens, G. W. (2017). The state of applied econometrics: Causality and policy evaluation. Journal of Economic Perspectives, 31(2), 3–32.

[3] Badia, A. P., Piot, B., Kapturowski, S., et al. (202

0). Agent57: Outperforming the Atari human benchmark. In Proceedings of the 37th International Conference on Machine Learning (pp. 507–517). PMLR.

[4] Baker, A., Dredze, M., & Osoba, O. A. (2022). A framework for responsible AI deployment in health systems. Journal of the American Medical Informatics Association, 29(12), 2069–2077.

[5] Balduzzi, D., Garnelo, M., Bachrach, Y., et al. (2019). Open-ended learning in symmetric zero-sum games. In Proceedings of the 36th International Conference on Machine Learning (pp. 434–443). PMLR.

[6] Berwick, D. M., Nolan, T. W., & Whittington, J. (2008). The triple aim: Care, health, and cost. Health Affairs, 27(3), 759–769.

[7] Bertsimas, D., Ling, D., & Schulman, L. (2020). Interpretable and parsimonious models for healthcare operations and policy. Management Science, 66(6), 2411–2432.

[8] Bica, I., Alaa, A. M., Lambert, C., & van der Schaar, M. (2020). From real-world patient data to individualized treatment effects using counterfactual inference. Advances in Neural Information Processing Systems, 33, 1146–1157.

[9] Bitterman, D. S., Aerts, H. J. W. L., Mak, R. H., & Bachireddy, P. (2020). AI in population health: Opportunities and pitfalls. JAMA, 323(6), 517–518.

[10] Boehmke, B., & Greenwell, B. (2019). Hands-on machine learning with R. CRC Press.

[11] Bolhuis, D., Wouters, O. J., & McKee, M. (2021). Value-based healthcare: An evidence-informed review. The Lancet Public Health, 6(10), e740–e748.

[12] Bousquet, J., Anto, J. M., Sterk, P. J., et al. (2021). Digital transformation of chronic disease management: Population health implications. The Lancet Digital Health, 3(6), e353–e362.

[13] Boutilier, C., et al. (2021). AI for healthcare: Challenges in deployment and governance. Nature Medicine, 27(11), 1812–1814.

[14] Bradley, E. H., Elkins, B. R., Herrin, J., & Elbel, B. (2018). Health and social services expenditures: Association with health outcomes. BMJ Quality & Safety, 27(10), 826–833.

[15] Braverman, A., Dai, A. M., Li, M., & Schalick, J. (2020). Deep learning for Medicare risk adjustment. Health Services Research, 55(1), 68–77.

[16] Broyles, S., & Buerhaus, P. I. (2020). Value-based payment models and outcomes: Implications for population health. Nursing Outlook, 68(5), 517–525.

[17] Buchman, T. G., & Simpson, S. Q. (2020). Digital twins in critical care and population health. Critical Care Medicine, 48(12), 1908–1910.

[18] Cassel, C. K., & Jain, S. H. (2021). Assessing value in healthcare: Measurement and incentives. New England Journal of Medicine, 384(7), 597–599.

[19] Chen, I. Y., Joshi, S., Ghassemi, M., & Ranganath, R. (2021). Probabilistic machine learning for healthcare. Annual Review of Biomedical Data Science, 4, 393–419.

[20] Cheng, Y., Zhao, W., Zhang, H., & Shen, D. (2022). Reinforcement learning for health: A survey. ACM Computing Surveys, 55(1), 1–36.

[21] Chowdhury, A., & Kuo, A. M. H. (2021). Population health analytics in the era of big data: A systematic review. Journal of Biomedical Informatics, 115, 103684.

[22] Chua, K.-P., & Conti, R. M. (2020). Value-based payment reforms: Effects and equity considerations. JAMA, 323(8), 707–708.

[23] Crites, S. L., & Buntin, M. B. (2020). The promise and risk of AI-driven payment and delivery reform. Health Affairs, 39(10), 1691–1697.

[24] Cutler, D. M., & Ghosh, K. (2012). The potential for cost savings through bundled episode payments. New England Journal of Medicine, 366(12), 1075–1077.

[25] Dai, A. M., Priskin, E., & Braverman, A. (2019). Improving risk adjustment with machine learning. Health Services Research, 54(S2), 1115–1124.

[26] Deaton, A., & Cartwright, N. (2018). Understanding and misunderstanding randomized controlled trials. Social Science & Medicine, 210, 2–21.

[27] Dumitrescu, A., van der Schaar, M., & Alaa, A. M. (2022). Off-policy evaluation for healthcare decision support: A review. Patterns, 3(7), 100520.

[28] Dulac-Arnold, G., Levine, N., Mankowitz, D. J., et al. (2020). An empirical investigation of the challenges of real-world reinforcement learning. arXiv.

[29] Dusenberry, M. W., et al. (2020). Analyzing the role of model uncertainty for electronic health record predictions. npj Digital Medicine, 3, 48.

[30] Economou, A., & Sokol-Hessner, P. (2022). Behavioral economics and incentives in value-based care. Health Affairs, 41(3), 344–351.

[31] Ellis, R. P., & McGuire, T. G. (2007). Predictability and predictiveness in health care spending. Journal of Health Economics, 26(1), 25–48.

[32] Emanuel, E. J., & Navathe, A. S. (2020). Will 30% of Medicare payments be in alternative payment models by 2030? JAMA, 323(9), 817–818.

[33] Erdem, E., & Tsoukalas, I. (2022). Digital twins for healthcare: A systematic review. IEEE Access, 10, 12388–12410.

[34] Esteva, A., Robicquet, A., Ramsundar, B., et al. (2019). A guide to deep learning in healthcare. Nature Medicine, 25(1), 24–29.

[35] Farias, V. F., & Li, L. (2022). Policy evaluation and optimization in contextual bandits. Foundations and Trends in Machine Learning, 15(2), 109–214.

[36] Ferreira, F., et al. (2021). Multi-agent simulation and reinforcement learning for health policy design: A review. Journal of Biomedical Informatics, 118, 103776.

[37] Finlayson, S. G., Bowers, J. D., Ito, J., et al. (2019). Adversarial attacks on medical machine learning. Science, 363(6433), 1287–1289.

[38] Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., & Whiteson, S. (2018). Counterfactual multi-agent policy gradients. In Proceedings of the AAAI Conference on Artificial Intelligence (pp. 2974–2982). AAAI Press.

[39] Fréchet, A., & Kober, J. (2020). A survey of deep reinforcement learning for multi-agent systems. Artificial Intelligence Review, 53(2), 1109–1148.

[40] Fudenberg, D., & Tirole, J. (1991). Game theory. MIT Press.

[41] Gartner, D. R., Padman, R., & Patel, N. (2020). Population health analytics: Methods and implementation. INFORMS Journal on Computing, 32(4), 934–950.

[42] Ghassemi, M., Oakden-Rayner, L., & Beam, A. L. (2021). The false hope of current approaches to explainable artificial intelligence in health care. The Lancet Digital Health, 3(11), e745–e750.

[43] Gottesman, O., Johansson, F., Komorowski, M., et al. (2019). Guidelines for reinforcement learning in healthcare. Nature Medicine, 25(1), 16–18.

[44] Greenhalgh, T., Wherton, J., Papoutsi, C., et al. (2017). Beyond adoption: A new framework for theorizing and evaluating nonadoption, abandonment, and challenges to scale-up of health technologies. Journal of Medical Internet Research, 19(11), e367.

[45] Gupta, A., et al. (2021). Off-policy evaluation in reinforcement learning for healthcare: A systematic review. IEEE Journal of Biomedical and Health Informatics, 25(12), 4226–4240.

[46] Hardt, M., Price, E., & Srebro, N. (2016). Equality of opportunity in supervised learning. Advances in Neural Information Processing Systems, 29, 3315–3323.

[47] Hearn, J., & Eggleston, K. (2022). Incentives and strategic behavior under value-based payment. Health Economics, 31(10), 2081–2097.

[48] Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., & Meger, D. (2018). Deep reinforcement learning that matters. In Proceedings of the AAAI Conference on Artificial Intelligence (pp. 3207–3214). AAAI Press.

[49] Himmelstein, D. U., & Woolhandler, S. (2020). The role of administrative costs in US health care. JAMA, 323(21), 2141–2142.

[50] Huang, T., et al. (2022). Digital twin-driven decision support for population health: A scoping review. The Lancet Digital Health, 4(8), e564–e575.

[51] Jain, S. H., & Cassel, C. K. (2020). Value-based care and measurement: What matters and what doesn’t. New England Journal of Medicine, 383(15), 1447–1449.

[52] Jiang, J., & Li, Y. (2021). Multi-agent reinforcement learning: A survey. ACM Computing Surveys, 54(9), 1–43.

[53] Johansson, F., Shalit, U., & Sontag, D. (2016). Learning representations for counterfactual inference. In Proceedings of the 33rd International Conference on Machine Learning (pp. 3020–3029). PMLR.

[54] Komorowski, M., Celi, L. A., Badawi, O., Gordon, A. C., & Faisal, A. A. (2018). The Artificial Intelligence Clinician learns optimal treatment strategies for sepsis in intensive care. Nature Medicine, 24(11), 1716–1720.

[55] Kruse, C. S., Beane, A., & Chokshi, D. A. (2022). Population health management and the digital transformation of care: A systematic review. Journal of Medical Internet Research, 24(6), e35220.

[56] Kwon, J.-M., Lee, Y., Lee, Y., Lee, S., & Park, J. (2018). An algorithm based on deep learning for predicting in-hospital cardiac arrest. The Lancet, 392(10148), 1577–1578.

[57] Langford, J., & Zhang, T. (2008). The epoch-greedy algorithm for contextual multi-armed bandits. Advances in Neural Information Processing Systems, 20, 817–824.

[58] LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444.

[59] Li, L., Chu, W., Langford, J., & Schapire, R. E. (2010). A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web (pp. 661–670). ACM.

[60] Lillicrap, T. P., Hunt, J. J., Pritzel, A., et al. (2016). Continuous control with deep reinforcement learning. In Proceedings of ICLR 2016.

[61] Liu, S. X., & Wang, Y. (2020). Value-based payment models: A review of evidence and implementation. The American Journal of Managed Care, 26(9), e280–e286.

[62] Ma, X., Kamar, E., & Horvitz, E. (2020). Multi-agent reinforcement learning for decision support: Foundations and human-in-the-loop considerations. AI Magazine, 41(2), 37–50.

[63] Mahmoudi, E., & Jensen, G. A. (2018). A review of risk adjustment in value-based payment. Medical Care Research and Review, 75(3), 323–348.

[64] Mak, H.-Y., Rong, Y., & Sun, J. (2015). Multi-agent reinforcement learning for dynamic contract design. Management Science, 61(10), 2380–2397.

[65] McGuire, T. G., Glazer, J., Newhouse, J. P., et al. (2013). Paying for performance in health insurance: Selection, incentives, and risk adjustment. Journal of Health Economics, 32(5), 901–912.

[66] Miikkulainen, R., et al. (2021). Evolving deep reinforcement learning policies at scale. Nature Machine Intelligence, 3(7), 588–595.

[67] Milstein, A., & Gilbertson, E. (2020). American medical home and population health outcomes under value-based payment. Health Affairs, 39(3), 384–391.

[68] Mitchell, M., Wu, S., Zaldivar, A., et al. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220–229). ACM.

[69] Mohammad, A., et al. (2023). Digital twins and simulation for healthcare operations and policy: A systematic review. Computers & Industrial Engineering, 176, 108965.

[70] Monroe, A., et al. (2021). Learning health systems and population health management: A scoping review. Journal of the American Medical Informatics Association, 28(9), 1956–1965.

[71] Navathe, A. S., Emanuel, E. J., Bond, A. M., et al. (2019). Association between the Medicare Shared Savings Program and cost outcomes. JAMA, 321(5), 1–11.

[72] Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453.

[73] Ostrovsky, A., & Barnett, M. L. (2020). Payment reform: Progress and pitfalls in alternative payment models. New England Journal of Medicine, 383(10), 903–905.

[74] Porter, M. E. (2010). What is value in health care? New England Journal of Medicine, 363(26), 2477–2481.

[75] Puterman, M. L. (2014). Markov decision processes: Discrete stochastic dynamic programming. Wiley.

[76] Rajkomar, A., Hardt, M., Howell, M. D., Corrado, G., & Chin, M. H. (2018). Ensuring fairness in machine learning to advance health equity. Annals of Internal Medicine, 169(12), 866–872.

[77] Rutherford, A., et al. (2022). Off-policy evaluation for sequential decision making in healthcare: Methods, assumptions, and challenges. Journal of Biomedical Informatics, 133, 104148.

[78] Saria, S., & Subbaswamy, A. (2019). Tutorial: Safe and reliable machine learning. arXiv.

[79] Schapire, R. E., & Freund, Y. (2012). Boosting: Foundations and algorithms. MIT Press.

[80] Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv.

[81] Shah, N. H., & Milstein, A. (2020). Bagley: Digital transformation and the economics of population health. Health Affairs, 39(2), 196–202.

[82] Shalit, U., Johansson, F. D., & Sontag, D. (2017). Estimating individual treatment effect: Generalization bounds and algorithms. In Proceedings of the 34th International Conference on Machine Learning (pp. 3076–3085). PMLR.

[83] Silver, D., Hubert, T., Schrittwieser, J., et al. (2018). A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science, 362(6419), 1140–1144.

[84] Song, H., et al. (2021). A digital twin for patient and population health: A systematic review of applications and enabling technologies. npj Digital Medicine, 4, 101.

[85] Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.

[86] Thaler, R. H., & Sunstein, C. R. (2008). Nudge: Improving decisions about health, wealth, and happiness. Yale University Press.

[87] Thornton, S., & Sinha, S. (2021). Principal–agent problems and incentive design in value-based payment. Journal of Health Economics, 78, 102482.

[88] van der Schaar, M., Alaa, A. M., Floto, A., et al. (2021). How artificial intelligence and machine learning can help healthcare systems respond to COVID-19. Machine Learning, 110(1), 1–12.

[89] Velu, A. V., et al. (2020). Population health management: A systematic review of interventions and outcomes. BMC Health Services Research, 20, 1032.

[90] Watkins, C. J. C. H., & Dayan, P. (1992). Q-learning. Machine Learning, 8(3–4), 279–292.

Additional Files

Published

2025-03-09

How to Cite

Multi-Agent RL for Value-Based Population Health Optimization. (2025). The American Online Journal of Science and Engineering (AOJSE), 3(01). https://doi.org/10.5281/zenodo.20744351

Similar Articles

1-10 of 23

You may also start an advanced similarity search for this article.