Multi-Agent RL for Value-Based Population Health Optimization
DOI:
https://doi.org/10.5281/zenodo.20744351Keywords:
Agentic Artificial Intelligence, Multi-Agent Reinforcement Learning (MARL), Population Health Management, Risk- and Performance-Based Contracts, Healthcare Payment Redistribution, Actor–Network Healthcare Models, Complex Adaptive Healthcare Systems, Agent-Based Decision Support, Health Information Systems Innovation, Risk-Adjusted Outcome Optimization, Healthcare Econometrics, Decentralized Healthcare Governance, AI-DrivenAbstract
Agentic artificial intelligence (AI) systems, including those realized through multi-agent reinforcement learning (MARL), are emerging in multiple domains, including information filtering, navigation, gaming, and robotics. These novel approaches may also be relevant to healthcare delivery and policy, as MARL is well suited to complex strategic problems that Richard Thaler characterized as “makeshift solutions to impossibly complex problems.” Population health management is such an arena, particularly in optimizing risk- and performance-based contracts, payment redistribution across agents and cadre types, and performance-related benchmarks for an actor-network consisting of several competing provider-organizations functioning in a shared ecosystem. Health information systems, including such agentic, MARL-based systems, could help realize “better, faster, cheaper, and ‘kinder’ healthcare” by improving risk-adjusted health and cost outcomes across multiple resident cohorts, thereby more effectively realizing the goals of modern econometrics in medicine and public health.
Current culture tends to represent healthcare decision-making as a Black Box or an informationally centralized “big brain” where supervision assures correctness and compliance with Thaler’s ideal of omniscience for coordinating activities. Yet actual healthcare organizations are not Black Boxes, much less a single Big Brain, corporate or otherwise. Decision makers are often overpromoted, misinformed, late in their decisions, veto the best ideas of the organization, or are unable to manage rich information effectively. Supporting such overloaded Decision Makers-MD, using training and expertise gained trying to wisely harvest large but incomplete information- is a hard problem. Because the decisions should be seen as Agents making active decisions, it is viewed in terms of Agent-based Systems, or the study of multi-oriented organizations as complex adaptive systems.
References
[1] Agarwal, A., Dudík, M., Langford, J., & Li, L. (2014). A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web (pp. 661–670). ACM.
[2] Athey, S., & Imbens, G. W. (2017). The state of applied econometrics: Causality and policy evaluation. Journal of Economic Perspectives, 31(2), 3–32.
[3] Badia, A. P., Piot, B., Kapturowski, S., et al. (202
0). Agent57: Outperforming the Atari human benchmark. In Proceedings of the 37th International Conference on Machine Learning (pp. 507–517). PMLR.
[4] Baker, A., Dredze, M., & Osoba, O. A. (2022). A framework for responsible AI deployment in health systems. Journal of the American Medical Informatics Association, 29(12), 2069–2077.
[5] Balduzzi, D., Garnelo, M., Bachrach, Y., et al. (2019). Open-ended learning in symmetric zero-sum games. In Proceedings of the 36th International Conference on Machine Learning (pp. 434–443). PMLR.
[6] Berwick, D. M., Nolan, T. W., & Whittington, J. (2008). The triple aim: Care, health, and cost. Health Affairs, 27(3), 759–769.
[7] Bertsimas, D., Ling, D., & Schulman, L. (2020). Interpretable and parsimonious models for healthcare operations and policy. Management Science, 66(6), 2411–2432.
[8] Bica, I., Alaa, A. M., Lambert, C., & van der Schaar, M. (2020). From real-world patient data to individualized treatment effects using counterfactual inference. Advances in Neural Information Processing Systems, 33, 1146–1157.
[9] Bitterman, D. S., Aerts, H. J. W. L., Mak, R. H., & Bachireddy, P. (2020). AI in population health: Opportunities and pitfalls. JAMA, 323(6), 517–518.
[10] Boehmke, B., & Greenwell, B. (2019). Hands-on machine learning with R. CRC Press.
[11] Bolhuis, D., Wouters, O. J., & McKee, M. (2021). Value-based healthcare: An evidence-informed review. The Lancet Public Health, 6(10), e740–e748.
[12] Bousquet, J., Anto, J. M., Sterk, P. J., et al. (2021). Digital transformation of chronic disease management: Population health implications. The Lancet Digital Health, 3(6), e353–e362.
[13] Boutilier, C., et al. (2021). AI for healthcare: Challenges in deployment and governance. Nature Medicine, 27(11), 1812–1814.
[14] Bradley, E. H., Elkins, B. R., Herrin, J., & Elbel, B. (2018). Health and social services expenditures: Association with health outcomes. BMJ Quality & Safety, 27(10), 826–833.
[15] Braverman, A., Dai, A. M., Li, M., & Schalick, J. (2020). Deep learning for Medicare risk adjustment. Health Services Research, 55(1), 68–77.
[16] Broyles, S., & Buerhaus, P. I. (2020). Value-based payment models and outcomes: Implications for population health. Nursing Outlook, 68(5), 517–525.
[17] Buchman, T. G., & Simpson, S. Q. (2020). Digital twins in critical care and population health. Critical Care Medicine, 48(12), 1908–1910.
[18] Cassel, C. K., & Jain, S. H. (2021). Assessing value in healthcare: Measurement and incentives. New England Journal of Medicine, 384(7), 597–599.
[19] Chen, I. Y., Joshi, S., Ghassemi, M., & Ranganath, R. (2021). Probabilistic machine learning for healthcare. Annual Review of Biomedical Data Science, 4, 393–419.
[20] Cheng, Y., Zhao, W., Zhang, H., & Shen, D. (2022). Reinforcement learning for health: A survey. ACM Computing Surveys, 55(1), 1–36.
[21] Chowdhury, A., & Kuo, A. M. H. (2021). Population health analytics in the era of big data: A systematic review. Journal of Biomedical Informatics, 115, 103684.
[22] Chua, K.-P., & Conti, R. M. (2020). Value-based payment reforms: Effects and equity considerations. JAMA, 323(8), 707–708.
[23] Crites, S. L., & Buntin, M. B. (2020). The promise and risk of AI-driven payment and delivery reform. Health Affairs, 39(10), 1691–1697.
[24] Cutler, D. M., & Ghosh, K. (2012). The potential for cost savings through bundled episode payments. New England Journal of Medicine, 366(12), 1075–1077.
[25] Dai, A. M., Priskin, E., & Braverman, A. (2019). Improving risk adjustment with machine learning. Health Services Research, 54(S2), 1115–1124.
[26] Deaton, A., & Cartwright, N. (2018). Understanding and misunderstanding randomized controlled trials. Social Science & Medicine, 210, 2–21.
[27] Dumitrescu, A., van der Schaar, M., & Alaa, A. M. (2022). Off-policy evaluation for healthcare decision support: A review. Patterns, 3(7), 100520.
[28] Dulac-Arnold, G., Levine, N., Mankowitz, D. J., et al. (2020). An empirical investigation of the challenges of real-world reinforcement learning. arXiv.
[29] Dusenberry, M. W., et al. (2020). Analyzing the role of model uncertainty for electronic health record predictions. npj Digital Medicine, 3, 48.
[30] Economou, A., & Sokol-Hessner, P. (2022). Behavioral economics and incentives in value-based care. Health Affairs, 41(3), 344–351.
[31] Ellis, R. P., & McGuire, T. G. (2007). Predictability and predictiveness in health care spending. Journal of Health Economics, 26(1), 25–48.
[32] Emanuel, E. J., & Navathe, A. S. (2020). Will 30% of Medicare payments be in alternative payment models by 2030? JAMA, 323(9), 817–818.
[33] Erdem, E., & Tsoukalas, I. (2022). Digital twins for healthcare: A systematic review. IEEE Access, 10, 12388–12410.
[34] Esteva, A., Robicquet, A., Ramsundar, B., et al. (2019). A guide to deep learning in healthcare. Nature Medicine, 25(1), 24–29.
[35] Farias, V. F., & Li, L. (2022). Policy evaluation and optimization in contextual bandits. Foundations and Trends in Machine Learning, 15(2), 109–214.
[36] Ferreira, F., et al. (2021). Multi-agent simulation and reinforcement learning for health policy design: A review. Journal of Biomedical Informatics, 118, 103776.
[37] Finlayson, S. G., Bowers, J. D., Ito, J., et al. (2019). Adversarial attacks on medical machine learning. Science, 363(6433), 1287–1289.
[38] Foerster, J., Farquhar, G., Afouras, T., Nardelli, N., & Whiteson, S. (2018). Counterfactual multi-agent policy gradients. In Proceedings of the AAAI Conference on Artificial Intelligence (pp. 2974–2982). AAAI Press.
[39] Fréchet, A., & Kober, J. (2020). A survey of deep reinforcement learning for multi-agent systems. Artificial Intelligence Review, 53(2), 1109–1148.
[40] Fudenberg, D., & Tirole, J. (1991). Game theory. MIT Press.
[41] Gartner, D. R., Padman, R., & Patel, N. (2020). Population health analytics: Methods and implementation. INFORMS Journal on Computing, 32(4), 934–950.
[42] Ghassemi, M., Oakden-Rayner, L., & Beam, A. L. (2021). The false hope of current approaches to explainable artificial intelligence in health care. The Lancet Digital Health, 3(11), e745–e750.
[43] Gottesman, O., Johansson, F., Komorowski, M., et al. (2019). Guidelines for reinforcement learning in healthcare. Nature Medicine, 25(1), 16–18.
[44] Greenhalgh, T., Wherton, J., Papoutsi, C., et al. (2017). Beyond adoption: A new framework for theorizing and evaluating nonadoption, abandonment, and challenges to scale-up of health technologies. Journal of Medical Internet Research, 19(11), e367.
[45] Gupta, A., et al. (2021). Off-policy evaluation in reinforcement learning for healthcare: A systematic review. IEEE Journal of Biomedical and Health Informatics, 25(12), 4226–4240.
[46] Hardt, M., Price, E., & Srebro, N. (2016). Equality of opportunity in supervised learning. Advances in Neural Information Processing Systems, 29, 3315–3323.
[47] Hearn, J., & Eggleston, K. (2022). Incentives and strategic behavior under value-based payment. Health Economics, 31(10), 2081–2097.
[48] Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., & Meger, D. (2018). Deep reinforcement learning that matters. In Proceedings of the AAAI Conference on Artificial Intelligence (pp. 3207–3214). AAAI Press.
[49] Himmelstein, D. U., & Woolhandler, S. (2020). The role of administrative costs in US health care. JAMA, 323(21), 2141–2142.
[50] Huang, T., et al. (2022). Digital twin-driven decision support for population health: A scoping review. The Lancet Digital Health, 4(8), e564–e575.
[51] Jain, S. H., & Cassel, C. K. (2020). Value-based care and measurement: What matters and what doesn’t. New England Journal of Medicine, 383(15), 1447–1449.
[52] Jiang, J., & Li, Y. (2021). Multi-agent reinforcement learning: A survey. ACM Computing Surveys, 54(9), 1–43.
[53] Johansson, F., Shalit, U., & Sontag, D. (2016). Learning representations for counterfactual inference. In Proceedings of the 33rd International Conference on Machine Learning (pp. 3020–3029). PMLR.
[54] Komorowski, M., Celi, L. A., Badawi, O., Gordon, A. C., & Faisal, A. A. (2018). The Artificial Intelligence Clinician learns optimal treatment strategies for sepsis in intensive care. Nature Medicine, 24(11), 1716–1720.
[55] Kruse, C. S., Beane, A., & Chokshi, D. A. (2022). Population health management and the digital transformation of care: A systematic review. Journal of Medical Internet Research, 24(6), e35220.
[56] Kwon, J.-M., Lee, Y., Lee, Y., Lee, S., & Park, J. (2018). An algorithm based on deep learning for predicting in-hospital cardiac arrest. The Lancet, 392(10148), 1577–1578.
[57] Langford, J., & Zhang, T. (2008). The epoch-greedy algorithm for contextual multi-armed bandits. Advances in Neural Information Processing Systems, 20, 817–824.
[58] LeCun, Y., Bengio, Y., & Hinton, G. (2015). Deep learning. Nature, 521(7553), 436–444.
[59] Li, L., Chu, W., Langford, J., & Schapire, R. E. (2010). A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web (pp. 661–670). ACM.
[60] Lillicrap, T. P., Hunt, J. J., Pritzel, A., et al. (2016). Continuous control with deep reinforcement learning. In Proceedings of ICLR 2016.
[61] Liu, S. X., & Wang, Y. (2020). Value-based payment models: A review of evidence and implementation. The American Journal of Managed Care, 26(9), e280–e286.
[62] Ma, X., Kamar, E., & Horvitz, E. (2020). Multi-agent reinforcement learning for decision support: Foundations and human-in-the-loop considerations. AI Magazine, 41(2), 37–50.
[63] Mahmoudi, E., & Jensen, G. A. (2018). A review of risk adjustment in value-based payment. Medical Care Research and Review, 75(3), 323–348.
[64] Mak, H.-Y., Rong, Y., & Sun, J. (2015). Multi-agent reinforcement learning for dynamic contract design. Management Science, 61(10), 2380–2397.
[65] McGuire, T. G., Glazer, J., Newhouse, J. P., et al. (2013). Paying for performance in health insurance: Selection, incentives, and risk adjustment. Journal of Health Economics, 32(5), 901–912.
[66] Miikkulainen, R., et al. (2021). Evolving deep reinforcement learning policies at scale. Nature Machine Intelligence, 3(7), 588–595.
[67] Milstein, A., & Gilbertson, E. (2020). American medical home and population health outcomes under value-based payment. Health Affairs, 39(3), 384–391.
[68] Mitchell, M., Wu, S., Zaldivar, A., et al. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220–229). ACM.
[69] Mohammad, A., et al. (2023). Digital twins and simulation for healthcare operations and policy: A systematic review. Computers & Industrial Engineering, 176, 108965.
[70] Monroe, A., et al. (2021). Learning health systems and population health management: A scoping review. Journal of the American Medical Informatics Association, 28(9), 1956–1965.
[71] Navathe, A. S., Emanuel, E. J., Bond, A. M., et al. (2019). Association between the Medicare Shared Savings Program and cost outcomes. JAMA, 321(5), 1–11.
[72] Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453.
[73] Ostrovsky, A., & Barnett, M. L. (2020). Payment reform: Progress and pitfalls in alternative payment models. New England Journal of Medicine, 383(10), 903–905.
[74] Porter, M. E. (2010). What is value in health care? New England Journal of Medicine, 363(26), 2477–2481.
[75] Puterman, M. L. (2014). Markov decision processes: Discrete stochastic dynamic programming. Wiley.
[76] Rajkomar, A., Hardt, M., Howell, M. D., Corrado, G., & Chin, M. H. (2018). Ensuring fairness in machine learning to advance health equity. Annals of Internal Medicine, 169(12), 866–872.
[77] Rutherford, A., et al. (2022). Off-policy evaluation for sequential decision making in healthcare: Methods, assumptions, and challenges. Journal of Biomedical Informatics, 133, 104148.
[78] Saria, S., & Subbaswamy, A. (2019). Tutorial: Safe and reliable machine learning. arXiv.
[79] Schapire, R. E., & Freund, Y. (2012). Boosting: Foundations and algorithms. MIT Press.
[80] Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv.
[81] Shah, N. H., & Milstein, A. (2020). Bagley: Digital transformation and the economics of population health. Health Affairs, 39(2), 196–202.
[82] Shalit, U., Johansson, F. D., & Sontag, D. (2017). Estimating individual treatment effect: Generalization bounds and algorithms. In Proceedings of the 34th International Conference on Machine Learning (pp. 3076–3085). PMLR.
[83] Silver, D., Hubert, T., Schrittwieser, J., et al. (2018). A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science, 362(6419), 1140–1144.
[84] Song, H., et al. (2021). A digital twin for patient and population health: A systematic review of applications and enabling technologies. npj Digital Medicine, 4, 101.
[85] Sutton, R. S., & Barto, A. G. (2018). Reinforcement learning: An introduction (2nd ed.). MIT Press.
[86] Thaler, R. H., & Sunstein, C. R. (2008). Nudge: Improving decisions about health, wealth, and happiness. Yale University Press.
[87] Thornton, S., & Sinha, S. (2021). Principal–agent problems and incentive design in value-based payment. Journal of Health Economics, 78, 102482.
[88] van der Schaar, M., Alaa, A. M., Floto, A., et al. (2021). How artificial intelligence and machine learning can help healthcare systems respond to COVID-19. Machine Learning, 110(1), 1–12.
[89] Velu, A. V., et al. (2020). Population health management: A systematic review of interventions and outcomes. BMC Health Services Research, 20, 1032.
[90] Watkins, C. J. C. H., & Dayan, P. (1992). Q-learning. Machine Learning, 8(3–4), 279–292.
Additional Files
Published
Issue
Section
License
The American Online Journal of Science and Engineering (AOJSE) is an open-access journal, and all published articles are made freely available to readers worldwide under the terms of the Creative Commons Attribution 4.0 International License (CC BY 4.0). This license governs the use, sharing, adaptation, and distribution of all content published in AOJSE and is designed to promote the broadest possible dissemination and reuse of scholarly research.
Creative Commons Attribution 4.0 International License (CC BY 4.0)
Under the Creative Commons Attribution 4.0 International License, any user is free to share, copy, and redistribute the published material in any medium or format, and to adapt, remix, transform, and build upon the material for any purpose, including commercial use, provided that appropriate credit is given to the original authors and source. The license cannot be revoked as long as the terms are followed.
When using or redistributing content published in AOJSE, users must provide proper attribution by citing the original authors' names, the article title, the journal name (AOJSE), the volume and issue number, the year of publication, and the DOI or URL of the article. Users must also indicate if any changes were made to the original work. Attribution must be provided in a reasonable manner but must not suggest that the authors or AOJSE endorse the user or their use of the material.
Author Rights and Copyright Retention
AOJSE respects and upholds the intellectual property rights of all contributing authors. Authors who publish in AOJSE retain full copyright over their work. By submitting a manuscript to AOJSE, authors grant the journal a non-exclusive, worldwide, royalty-free license to publish, reproduce, distribute, publicly display, and archive the article in print and electronic formats. This license allows AOJSE to make the article freely accessible to readers globally while ensuring that the authors remain the rightful owners of their work.
Authors are permitted to deposit their published articles in institutional repositories, personal websites, academic networking platforms, and preprint servers, provided that the original publication in AOJSE is acknowledged and properly cited. Authors may also reuse their published content in subsequent works, presentations, teaching materials, and grant applications without requiring prior permission from the journal.
Third-Party Content
Authors are solely responsible for obtaining written permission to reproduce any third-party material, including figures, tables, images, data, or excerpts from other publications, included in their manuscript. Evidence of such permissions must be provided to the editorial office upon request. AOJSE does not assume any responsibility for copyright infringement arising from the unauthorized inclusion of third-party content in published articles.
Permitted Uses
Under the CC BY 4.0 license, the following uses of AOJSE published content are freely permitted without prior written permission from the journal or authors, provided that proper attribution is given. Users may read, download, print, and distribute articles for personal, educational, or research purposes. Researchers may reuse data, figures, and findings for meta-analyses, systematic reviews, and secondary research. Educators may incorporate published articles into course materials, syllabi, and academic presentations. Journalists, policymakers, and practitioners may reference and cite published findings for professional and public interest purposes.
Prohibited Uses
While the CC BY 4.0 license is broad and permissive, users may not apply legal terms or technological measures that legally restrict others from doing anything the license permits. Users may not misrepresent the original authorship of published content or imply endorsement by the original authors or AOJSE without explicit consent. Plagiarism, data fabrication, and any form of academic misconduct in the reuse of published material are strictly prohibited and are a violation of publication ethics standards.
Institutional and Repository Deposit Policy
AOJSE fully supports self-archiving and green open access. Authors are encouraged to deposit the final published version of their article, also known as the version of record, in institutional repositories, subject repositories, and open-access databases immediately upon publication. There is no embargo period. Authors should always link back to the original article on the AOJSE website and include the full citation and DOI when depositing their work in any repository.
Digital Preservation and Archiving
AOJSE is committed to the long-term digital preservation of all published content to ensure its permanent availability and accessibility. The journal utilizes the Open Journal Systems (OJS) platform, which supports internationally recognized digital archiving standards. AOJSE encourages participation in archiving networks such as LOCKSS (Lots of Copies Keep Stuff Safe) and CLOCKSS (Controlled LOCKSS) to safeguard published articles against data loss and ensure continued access for future generations of researchers and readers.
Disclaimer
The views, opinions, findings, and conclusions expressed in articles published in AOJSE are solely those of the authors and do not necessarily reflect the official position, policy, or views of the journal, its editorial board, or its affiliated organizations. AOJSE makes every effort to ensure the accuracy and integrity of published content but does not accept legal responsibility for any errors, omissions, or claims arising from the use of information published in the journal.
Contact for Licensing Inquiries
For any questions regarding the licensing terms, copyright, or permitted use of content published in AOJSE, please contact the editorial office through the journal website at https://aojse.org. The editorial team will be happy to assist authors, readers, and institutions with any licensing-related queries.