{"id":819473,"date":"2026-08-20T03:40:20","date_gmt":"2026-08-20T03:40:20","guid":{"rendered":"https:\/\/www.newsbeep.com\/us\/819473\/"},"modified":"2026-08-20T03:40:20","modified_gmt":"2026-08-20T03:40:20","slug":"safety-and-security-of-large-language-models-in-healthcare","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/us\/819473\/","title":{"rendered":"Safety and security of large language models in healthcare"},"content":{"rendered":"<p class=\"c-article-references__text\" id=\"ref-CR1\">Gommers, J. et al. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial. Lancet 407, 505\u2013514 (2026).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/S0140-6736(25)02464-X\" data-track-item_id=\"10.1016\/S0140-6736(25)02464-X\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2FS0140-6736%2825%2902464-X\" aria-label=\"Article reference 1\" data-doi=\"10.1016\/S0140-6736(25)02464-X\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=41620232\" aria-label=\"PubMed reference 1\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 1\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Interval%20cancer%2C%20sensitivity%2C%20and%20specificity%20comparing%20AI-supported%20mammography%20screening%20with%20standard%20double%20reading%20without%20AI%20in%20the%20MASAI%20study%3A%20a%20randomised%2C%20controlled%2C%20non-inferiority%2C%20single-blinded%2C%20population-based%2C%20screening-accuracy%20trial&amp;journal=Lancet&amp;doi=10.1016%2FS0140-6736%2825%2902464-X&amp;volume=407&amp;pages=505-514&amp;publication_year=2026&amp;author=Gommers%2CJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR2\">Lu, M. Y. et al. A multimodal generative AI copilot for human pathology. Nature 643, 466\u2013473 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41586-024-07618-3\" data-track-item_id=\"10.1038\/s41586-024-07618-3\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41586-024-07618-3\" aria-label=\"Article reference 2\" data-doi=\"10.1038\/s41586-024-07618-3\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2024Natur.634..466L\" aria-label=\"ADS reference 2\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2cXitVWgu7nK\" aria-label=\"CAS reference 2\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=38866050\" aria-label=\"PubMed reference 2\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11464372\" aria-label=\"PubMed Central reference 2\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 2\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=A%20multimodal%20generative%20AI%20copilot%20for%20human%20pathology&amp;journal=Nature&amp;doi=10.1038%2Fs41586-024-07618-3&amp;volume=634&amp;pages=466-473&amp;publication_year=2024&amp;author=Lu%2CMY\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR3\">Singh, R., Bapna, M., Diab, A. R., Ruiz, E. S. &amp; Lotter, W. How AI is used in FDA-authorized medical devices: a taxonomy across 1,016 authorizations. NPJ Digit. Med. 8, 388 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41746-025-01800-1\" data-track-item_id=\"10.1038\/s41746-025-01800-1\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41746-025-01800-1\" aria-label=\"Article reference 3\" data-doi=\"10.1038\/s41746-025-01800-1\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40596700\" aria-label=\"PubMed reference 3\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12219150\" aria-label=\"PubMed Central reference 3\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 3\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=How%20AI%20is%20used%20in%20FDA-authorized%20medical%20devices%3A%20a%20taxonomy%20across%201%2C016%20authorizations&amp;journal=NPJ%20Digit.%20Med.&amp;doi=10.1038%2Fs41746-025-01800-1&amp;volume=8&amp;publication_year=2025&amp;author=Singh%2CR&amp;author=Bapna%2CM&amp;author=Diab%2CAR&amp;author=Ruiz%2CES&amp;author=Lotter%2CW\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR4\">Griot, M., Vanderdonckt, J. &amp; Yuksel, D. Implementation of large language models in electronic health records. PLOS Digital Health <a href=\"https:\/\/doi.org\/10.1371\/journal.pdig.0001141\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.1371\/journal.pdig.0001141\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.1371\/journal.pdig.0001141<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR5\">Armitage, H. Clinicians can \u2018chat\u2019 with medical records through new AI software, ChatEHR. Stanford Medicine <a href=\"https:\/\/med.stanford.edu\/news\/all-news\/2025\/06\/chatehr.html\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/med.stanford.edu\/news\/all-news\/2025\/06\/chatehr.html\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/med.stanford.edu\/news\/all-news\/2025\/06\/chatehr.html<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR6\">Moor, M. et al. Foundation models for generalist medical artificial intelligence. Nature 616, 259\u2013265 (2023).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41586-023-05881-4\" data-track-item_id=\"10.1038\/s41586-023-05881-4\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41586-023-05881-4\" aria-label=\"Article reference 6\" data-doi=\"10.1038\/s41586-023-05881-4\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2023Natur.616..259M\" aria-label=\"ADS reference 6\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB3sXns1yqu70%3D\" aria-label=\"CAS reference 6\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=37045921\" aria-label=\"PubMed reference 6\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 6\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Foundation%20models%20for%20generalist%20medical%20artificial%20intelligence&amp;journal=Nature&amp;doi=10.1038%2Fs41586-023-05881-4&amp;volume=616&amp;pages=259-265&amp;publication_year=2023&amp;author=Moor%2CM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR7\">Clusmann, J. et al. The future landscape of large language models in medicine. Commun. Med. 3, 141 (2023).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s43856-023-00370-1\" data-track-item_id=\"10.1038\/s43856-023-00370-1\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs43856-023-00370-1\" aria-label=\"Article reference 7\" data-doi=\"10.1038\/s43856-023-00370-1\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=37816837\" aria-label=\"PubMed reference 7\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC10564921\" aria-label=\"PubMed Central reference 7\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 7\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=The%20future%20landscape%20of%20large%20language%20models%20in%20medicine&amp;journal=Commun.%20Med.&amp;doi=10.1038%2Fs43856-023-00370-1&amp;volume=3&amp;publication_year=2023&amp;author=Clusmann%2CJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR8\">Bubeck, S. et al. Sparks of Artificial General Intelligence: Early experiments with GPT-4. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2303.12712\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2303.12712\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2303.12712<\/a> (2023).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR9\">Singhal, K. et al. Toward expert-level medical question answering with large language models. Nat. Med. 31, 943\u2013950 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41591-024-03423-7\" data-track-item_id=\"10.1038\/s41591-024-03423-7\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41591-024-03423-7\" aria-label=\"Article reference 9\" data-doi=\"10.1038\/s41591-024-03423-7\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXhtVWru7s%3D\" aria-label=\"CAS reference 9\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39779926\" aria-label=\"PubMed reference 9\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11922739\" aria-label=\"PubMed Central reference 9\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 9\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Toward%20expert-level%20medical%20question%20answering%20with%20large%20language%20models&amp;journal=Nat.%20Med.&amp;doi=10.1038%2Fs41591-024-03423-7&amp;volume=31&amp;pages=943-950&amp;publication_year=2025&amp;author=Singhal%2CK\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR10\">Gu, Y. et al. The illusion of readiness: stress testing large frontier models on multimodal medical benchmarks. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2509.18234\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2509.18234\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2509.18234<\/a> (2025). This study reveals that medical LLM benchmarks may overestimate real-world readiness as benchmarks fail to capture brittleness and reasoning flaws.<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR11\">Omar, M. et al. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Commun. Med. 5, 330 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s43856-025-01021-3\" data-track-item_id=\"10.1038\/s43856-025-01021-3\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs43856-025-01021-3\" aria-label=\"Article reference 11\" data-doi=\"10.1038\/s43856-025-01021-3\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40753316\" aria-label=\"PubMed reference 11\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12318031\" aria-label=\"PubMed Central reference 11\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 11\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Multi-model%20assurance%20analysis%20showing%20large%20language%20models%20are%20highly%20vulnerable%20to%20adversarial%20hallucination%20attacks%20during%20clinical%20decision%20support&amp;journal=Commun.%20Med.&amp;doi=10.1038%2Fs43856-025-01021-3&amp;volume=5&amp;publication_year=2025&amp;author=Omar%2CM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR12\">Bedi, S., Jiang, Y., Chung, P., Koyejo, S. &amp; Shah, N. Fidelity of medical reasoning in large language models. JAMA Netw. Open 8, e2526021 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1001\/jamanetworkopen.2025.26021\" data-track-item_id=\"10.1001\/jamanetworkopen.2025.26021\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1001%2Fjamanetworkopen.2025.26021\" aria-label=\"Article reference 12\" data-doi=\"10.1001\/jamanetworkopen.2025.26021\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40779272\" aria-label=\"PubMed reference 12\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12334947\" aria-label=\"PubMed Central reference 12\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 12\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Fidelity%20of%20medical%20reasoning%20in%20large%20language%20models&amp;journal=JAMA%20Netw.%20Open&amp;doi=10.1001%2Fjamanetworkopen.2025.26021&amp;volume=8&amp;publication_year=2025&amp;author=Bedi%2CS&amp;author=Jiang%2CY&amp;author=Chung%2CP&amp;author=Koyejo%2CS&amp;author=Shah%2CN\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR13\">Goh, E. et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nat. Med. 31, 1233\u20131238 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41591-024-03456-y\" data-track-item_id=\"10.1038\/s41591-024-03456-y\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41591-024-03456-y\" aria-label=\"Article reference 13\" data-doi=\"10.1038\/s41591-024-03456-y\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXjtVKjtL8%3D\" aria-label=\"CAS reference 13\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39910272\" aria-label=\"PubMed reference 13\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12380382\" aria-label=\"PubMed Central reference 13\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 13\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=GPT-4%20assistance%20for%20improvement%20of%20physician%20performance%20on%20patient%20care%20tasks%3A%20a%20randomized%20controlled%20trial&amp;journal=Nat.%20Med.&amp;doi=10.1038%2Fs41591-024-03456-y&amp;volume=31&amp;pages=1233-1238&amp;publication_year=2025&amp;author=Goh%2CE\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR14\">McDuff, D. Towards accurate differential diagnosis with large language models. Nature 642, 451\u2013457 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41586-025-08869-4\" data-track-item_id=\"10.1038\/s41586-025-08869-4\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41586-025-08869-4\" aria-label=\"Article reference 14\" data-doi=\"10.1038\/s41586-025-08869-4\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2025Natur.642..451M\" aria-label=\"ADS reference 14\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXhtlGis73J\" aria-label=\"CAS reference 14\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40205049\" aria-label=\"PubMed reference 14\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12158753\" aria-label=\"PubMed Central reference 14\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 14\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Towards%20accurate%20differential%20diagnosis%20with%20large%20language%20models&amp;journal=Nature&amp;doi=10.1038%2Fs41586-025-08869-4&amp;volume=642&amp;pages=451-457&amp;publication_year=2025&amp;author=McDuff%2CD\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR15\">Tu, T. et al. Towards conversational diagnostic artificial intelligence. Nature 642, 442\u2013450 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41586-025-08866-7\" data-track-item_id=\"10.1038\/s41586-025-08866-7\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41586-025-08866-7\" aria-label=\"Article reference 15\" data-doi=\"10.1038\/s41586-025-08866-7\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2025Natur.642..442T\" aria-label=\"ADS reference 15\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXhtlGis73E\" aria-label=\"CAS reference 15\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40205050\" aria-label=\"PubMed reference 15\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12158756\" aria-label=\"PubMed Central reference 15\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 15\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Towards%20conversational%20diagnostic%20artificial%20intelligence&amp;journal=Nature&amp;doi=10.1038%2Fs41586-025-08866-7&amp;volume=642&amp;pages=442-450&amp;publication_year=2025&amp;author=Tu%2CT\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR16\">Palepu, A., et al. Exploring large language models for specialist-level oncology care. NEJM AI 2, AIcs2500025 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1056\/AIcs2500025\" data-track-item_id=\"10.1056\/AIcs2500025\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1056%2FAIcs2500025\" aria-label=\"Article reference 16\" data-doi=\"10.1056\/AIcs2500025\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 16\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Exploring%20large%20language%20models%20for%20specialist-level%20oncology%20care&amp;journal=NEJM%20AI&amp;doi=10.1056%2FAIcs2500025&amp;volume=2&amp;publication_year=2025&amp;author=Palepu%2CA\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR17\">Singhal, K. et al. Large language models encode clinical knowledge. Nature 620, 172\u2013180 (2023).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41586-023-06291-2\" data-track-item_id=\"10.1038\/s41586-023-06291-2\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41586-023-06291-2\" aria-label=\"Article reference 17\" data-doi=\"10.1038\/s41586-023-06291-2\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2023Natur.620..172S\" aria-label=\"ADS reference 17\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB3sXhsVKju7zP\" aria-label=\"CAS reference 17\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=37438534\" aria-label=\"PubMed reference 17\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC10396962\" aria-label=\"PubMed Central reference 17\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 17\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Large%20language%20models%20encode%20clinical%20knowledge&amp;journal=Nature&amp;doi=10.1038%2Fs41586-023-06291-2&amp;volume=620&amp;pages=172-180&amp;publication_year=2023&amp;author=Singhal%2CK\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR18\">Thirunavukarasu, A. J. et al. Large language models in medicine. Nat. Med. 29, 1930\u20131940 (2023).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41591-023-02448-8\" data-track-item_id=\"10.1038\/s41591-023-02448-8\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41591-023-02448-8\" aria-label=\"Article reference 18\" data-doi=\"10.1038\/s41591-023-02448-8\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB3sXhsVymtbbL\" aria-label=\"CAS reference 18\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=37460753\" aria-label=\"PubMed reference 18\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 18\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Large%20language%20models%20in%20medicine&amp;journal=Nat.%20Med.&amp;doi=10.1038%2Fs41591-023-02448-8&amp;volume=29&amp;pages=1930-1940&amp;publication_year=2023&amp;author=Thirunavukarasu%2CAJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR19\">Van Veen, D. et al. Adapted large language models can outperform medical experts in clinical text summarization. Nat. Med. 30, 1134\u20131142 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41591-024-02855-5\" data-track-item_id=\"10.1038\/s41591-024-02855-5\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41591-024-02855-5\" aria-label=\"Article reference 19\" data-doi=\"10.1038\/s41591-024-02855-5\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=38413730\" aria-label=\"PubMed reference 19\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11479659\" aria-label=\"PubMed Central reference 19\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 19\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Adapted%20large%20language%20models%20can%20outperform%20medical%20experts%20in%20clinical%20text%20summarization&amp;journal=Nat.%20Med.&amp;doi=10.1038%2Fs41591-024-02855-5&amp;volume=30&amp;pages=1134-1142&amp;publication_year=2024&amp;author=Veen%2CD\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR20\">Tierney, A. A. et al. Ambient artificial intelligence scribes: learnings after 1 year and over 2.5 million uses. NEJM Catal. Innov. Care Deliv. <a href=\"https:\/\/doi.org\/10.1056\/CAT.25.0040\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.1056\/CAT.25.0040\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.1056\/CAT.25.0040<\/a> (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1056\/CAT.25.0040\" data-track-item_id=\"10.1056\/CAT.25.0040\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1056%2FCAT.25.0040\" aria-label=\"Article reference 20\" data-doi=\"10.1056\/CAT.25.0040\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 20\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Ambient%20artificial%20intelligence%20scribes%3A%20learnings%20after%201%20year%20and%20over%202.5%20million%20uses&amp;journal=NEJM%20Catal.%20Innov.%20Care%20Deliv.&amp;doi=10.1056%2FCAT.25.0040&amp;publication_year=2025&amp;author=Tierney%2CAA\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR21\">Heinz, M. V. et al. Randomized trial of a generative AI chatbot for mental health treatment. NEJM AI 2, AIoa2400802 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 21\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Randomized%20trial%20of%20a%20generative%20AI%20chatbot%20for%20mental%20health%20treatment&amp;journal=NEJM%20AI&amp;volume=2&amp;publication_year=2025&amp;author=Heinz%2CMV\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR22\">Comanici, G. et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2507.06261\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2507.06261\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2507.06261<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR23\">Zakka, C. et al. Almanac \u2014 Retrieval-Augmented Language Models for Clinical Medicine. NEJM AI 1, AIoa2300068 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1056\/AIoa2300068\" data-track-item_id=\"10.1056\/AIoa2300068\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1056%2FAIoa2300068\" aria-label=\"Article reference 23\" data-doi=\"10.1056\/AIoa2300068\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 23\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Almanac%20%E2%80%94%20Retrieval-Augmented%20Language%20Models%20for%20Clinical%20Medicine&amp;journal=NEJM%20AI&amp;doi=10.1056%2FAIoa2300068&amp;volume=1&amp;publication_year=2024&amp;author=Zakka%2CC\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR24\">Wiest, I. C. et al. Large language models for clinical decision support in gastroenterology and hepatology. Nat. Rev. Gastroeneterol. Hepatol. 22, 773\u2013787 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41575-025-01108-1\" data-track-item_id=\"10.1038\/s41575-025-01108-1\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41575-025-01108-1\" aria-label=\"Article reference 24\" data-doi=\"10.1038\/s41575-025-01108-1\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 24\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Large%20language%20models%20for%20clinical%20decision%20support%20in%20gastroenterology%20and%20hepatology&amp;journal=Nat.%20Rev.%20Gastroeneterol.%20Hepatol.&amp;doi=10.1038%2Fs41575-025-01108-1&amp;volume=22&amp;pages=773-787&amp;publication_year=2025&amp;author=Wiest%2CIC\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR25\">Ferber, D. et al. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology. Nat. Cancer 6, 1337\u20131349 (2025). This study presents one of the first modular LLM-agent frameworks for healthcare, integrating biomedical tools and knowledge bases for clinical decision-making.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s43018-025-00991-6\" data-track-item_id=\"10.1038\/s43018-025-00991-6\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs43018-025-00991-6\" aria-label=\"Article reference 25\" data-doi=\"10.1038\/s43018-025-00991-6\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXhvVeqt7fI\" aria-label=\"CAS reference 25\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40481323\" aria-label=\"PubMed reference 25\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12380607\" aria-label=\"PubMed Central reference 25\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 25\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Development%20and%20validation%20of%20an%20autonomous%20artificial%20intelligence%20agent%20for%20clinical%20decision-making%20in%20oncology&amp;journal=Nat.%20Cancer&amp;doi=10.1038%2Fs43018-025-00991-6&amp;volume=6&amp;pages=1337-1349&amp;publication_year=2025&amp;author=Ferber%2CD\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR26\">Ge, J. et al. Development of a liver disease-specific large language model chat interface using retrieval-augmented generation. Hepatology 80, 1158\u20131168 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1097\/HEP.0000000000000834\" data-track-item_id=\"10.1097\/HEP.0000000000000834\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1097%2FHEP.0000000000000834\" aria-label=\"Article reference 26\" data-doi=\"10.1097\/HEP.0000000000000834\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=38451962\" aria-label=\"PubMed reference 26\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11706764\" aria-label=\"PubMed Central reference 26\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 26\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Development%20of%20a%20liver%20disease-specific%20large%20language%20model%20chat%20interface%20using%20retrieval-augmented%20generation&amp;journal=Hepatology&amp;doi=10.1097%2FHEP.0000000000000834&amp;volume=80&amp;pages=1158-1168&amp;publication_year=2024&amp;author=Ge%2CJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR27\">Kresevic, S. et al. Optimization of hepatological clinical guidelines interpretation by large language models: a retrieval augmented generation-based framework. NPJ Digit. Med. 7, 102 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41746-024-01091-y\" data-track-item_id=\"10.1038\/s41746-024-01091-y\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41746-024-01091-y\" aria-label=\"Article reference 27\" data-doi=\"10.1038\/s41746-024-01091-y\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=38654102\" aria-label=\"PubMed reference 27\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11039454\" aria-label=\"PubMed Central reference 27\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 27\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Optimization%20of%20hepatological%20clinical%20guidelines%20interpretation%20by%20large%20language%20models%3A%20a%20retrieval%20augmented%20generation-based%20framework&amp;journal=NPJ%20Digit.%20Med.&amp;doi=10.1038%2Fs41746-024-01091-y&amp;volume=7&amp;publication_year=2024&amp;author=Kresevic%2CS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR28\">Masanneck, L., Meuth, S. G. &amp; Pawlitzki, M. Evaluating base and retrieval augmented LLMs with document or online support for evidence based neurology. NPJ Digit. Med. 8, 137 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41746-025-01536-y\" data-track-item_id=\"10.1038\/s41746-025-01536-y\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41746-025-01536-y\" aria-label=\"Article reference 28\" data-doi=\"10.1038\/s41746-025-01536-y\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40038423\" aria-label=\"PubMed reference 28\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11880332\" aria-label=\"PubMed Central reference 28\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 28\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Evaluating%20base%20and%20retrieval%20augmented%20LLMs%20with%20document%20or%20online%20support%20for%20evidence%20based%20neurology&amp;journal=NPJ%20Digit.%20Med.&amp;doi=10.1038%2Fs41746-025-01536-y&amp;volume=8&amp;publication_year=2025&amp;author=Masanneck%2CL&amp;author=Meuth%2CSG&amp;author=Pawlitzki%2CM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR29\">Zakka, C. et al. Almanac Copilot: towards autonomous electronic health record navigation. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2405.07896\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2405.07896\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2405.07896<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR30\">Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 633\u2013638 (2025). This study shows that reinforcement learning can markedly improve LLM reasoning performance and introduces DeepSeek-R1 as a major open-weight model and one of few LLMs in peer-reviewed literature.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41586-025-09422-z\" data-track-item_id=\"10.1038\/s41586-025-09422-z\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41586-025-09422-z\" aria-label=\"Article reference 30\" data-doi=\"10.1038\/s41586-025-09422-z\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2025Natur.645..633G\" aria-label=\"ADS reference 30\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXitlSgt7rE\" aria-label=\"CAS reference 30\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40962978\" aria-label=\"PubMed reference 30\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12443585\" aria-label=\"PubMed Central reference 30\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 30\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=DeepSeek-R1%20incentivizes%20reasoning%20in%20LLMs%20through%20reinforcement%20learning&amp;journal=Nature&amp;doi=10.1038%2Fs41586-025-09422-z&amp;volume=645&amp;pages=633-638&amp;publication_year=2025&amp;author=Guo%2CD\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR31\">Angus, D. C. et al. AI, health, and health care today and tomorrow: the JAMA Summit report on artificial intelligence. JAMA 334, 1650\u20131664 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1001\/jama.2025.18490\" data-track-item_id=\"10.1001\/jama.2025.18490\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1001%2Fjama.2025.18490\" aria-label=\"Article reference 31\" data-doi=\"10.1001\/jama.2025.18490\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=41082366\" aria-label=\"PubMed reference 31\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 31\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=AI%2C%20health%2C%20and%20health%20care%20today%20and%20tomorrow%3A%20the%20JAMA%20Summit%20report%20on%20artificial%20intelligence&amp;journal=JAMA&amp;doi=10.1001%2Fjama.2025.18490&amp;volume=334&amp;pages=1650-1664&amp;publication_year=2025&amp;author=Angus%2CDC\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR32\">Garc\u00eda-Garc\u00eda, D., Le\u00f3n-G\u00f3mez, I., P\u00e9rez-Mar\u00edn, L. &amp; G\u00f3mez-Barroso, D. Exploring all-cause mortality surveillance during the Iberian Peninsula power outage, Spain, 28 April 2025. Eurosurveillance 30, 2500405 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.2807\/1560-7917.ES.2025.30.26.2500405\" data-track-item_id=\"10.2807\/1560-7917.ES.2025.30.26.2500405\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.2807%2F1560-7917.ES.2025.30.26.2500405\" aria-label=\"Article reference 32\" data-doi=\"10.2807\/1560-7917.ES.2025.30.26.2500405\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40613127\" aria-label=\"PubMed reference 32\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12231376\" aria-label=\"PubMed Central reference 32\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 32\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Exploring%20all-cause%20mortality%20surveillance%20during%20the%20Iberian%20Peninsula%20power%20outage%2C%20Spain%2C%2028%20April%202025&amp;journal=Eurosurveillance&amp;doi=10.2807%2F1560-7917.ES.2025.30.26.2500405&amp;volume=30&amp;publication_year=2025&amp;author=Garc%C3%ADa-Garc%C3%ADa%2CD&amp;author=Le%C3%B3n-G%C3%B3mez%2CI&amp;author=P%C3%A9rez-Mar%C3%ADn%2CL&amp;author=G%C3%B3mez-Barroso%2CD\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR33\">Larsen, E., Fong, A., Wernz, C. &amp; Ratwani, R. M. Implications of electronic health record downtime: an analysis of patient safety event reports. J. Am. Med. Inform. Assoc. 25, 187\u2013191 (2018).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1093\/jamia\/ocx057\" data-track-item_id=\"10.1093\/jamia\/ocx057\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1093%2Fjamia%2Focx057\" aria-label=\"Article reference 33\" data-doi=\"10.1093\/jamia\/ocx057\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=28575417\" aria-label=\"PubMed reference 33\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC7647128\" aria-label=\"PubMed Central reference 33\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 33\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Implications%20of%20electronic%20health%20record%20downtime%3A%20an%20analysis%20of%20patient%20safety%20event%20reports&amp;journal=J.%20Am.%20Med.%20Inform.%20Assoc.&amp;doi=10.1093%2Fjamia%2Focx057&amp;volume=25&amp;pages=187-191&amp;publication_year=2018&amp;author=Larsen%2CE&amp;author=Fong%2CA&amp;author=Wernz%2CC&amp;author=Ratwani%2CRM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR34\">Martin, G., Ghafur, S., Kinross, J., Hankin, C. &amp; Darzi, A. WannaCry\u2014a year on. BMJ 361, k2381 (2018).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1136\/bmj.k2381\" data-track-item_id=\"10.1136\/bmj.k2381\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1136%2Fbmj.k2381\" aria-label=\"Article reference 34\" data-doi=\"10.1136\/bmj.k2381\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=29866711\" aria-label=\"PubMed reference 34\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 34\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=WannaCry%E2%80%94a%20year%20on&amp;journal=BMJ&amp;doi=10.1136%2Fbmj.k2381&amp;volume=361&amp;publication_year=2018&amp;author=Martin%2CG&amp;author=Ghafur%2CS&amp;author=Kinross%2CJ&amp;author=Hankin%2CC&amp;author=Darzi%2CA\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR35\">Law, R. Cyberattacks on healthcare: Russia\u2019s tool for mass disruption. Medical Device Network <a href=\"https:\/\/www.medicaldevice-network.com\/features\/cyberattacks-on-healthcare-russias-tool-for-mass-disruption\/\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/www.medicaldevice-network.com\/features\/cyberattacks-on-healthcare-russias-tool-for-mass-disruption\/\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.medicaldevice-network.com\/features\/cyberattacks-on-healthcare-russias-tool-for-mass-disruption\/<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR36\">Cartwright, A. J. The elephant in the room: cybersecurity in healthcare. J. Clin. Monit. Comput. 37, 1123\u20131132 (2023).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"noopener nofollow\" data-track-label=\"10.1007\/s10877-023-01013-5\" data-track-item_id=\"10.1007\/s10877-023-01013-5\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/link.springer.com\/doi\/10.1007\/s10877-023-01013-5\" aria-label=\"Article reference 36\" data-doi=\"10.1007\/s10877-023-01013-5\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=37088852\" aria-label=\"PubMed reference 36\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC10123010\" aria-label=\"PubMed Central reference 36\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 36\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=The%20elephant%20in%20the%20room%3A%20cybersecurity%20in%20healthcare&amp;journal=J.%20Clin.%20Monit.%20Comput.&amp;doi=10.1007%2Fs10877-023-01013-5&amp;volume=37&amp;pages=1123-1132&amp;publication_year=2023&amp;author=Cartwright%2CAJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR37\">Gordon, W. J. et al. Assessment of employee susceptibility to phishing attacks at US health care institutions. JAMA Netw. Open 2, e190393 (2019).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1001\/jamanetworkopen.2019.0393\" data-track-item_id=\"10.1001\/jamanetworkopen.2019.0393\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1001%2Fjamanetworkopen.2019.0393\" aria-label=\"Article reference 37\" data-doi=\"10.1001\/jamanetworkopen.2019.0393\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=30848810\" aria-label=\"PubMed reference 37\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC6484661\" aria-label=\"PubMed Central reference 37\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 37\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Assessment%20of%20employee%20susceptibility%20to%20phishing%20attacks%20at%20US%20health%20care%20institutions&amp;journal=JAMA%20Netw.%20Open&amp;doi=10.1001%2Fjamanetworkopen.2019.0393&amp;volume=2&amp;publication_year=2019&amp;author=Gordon%2CWJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR38\">Li, S., Surineni, K. &amp; Prabhakaran, N. Cyber-attacks on hospital systems: a narrative review. Am. J. Ger. Psychiatry 7, 30\u201339 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 38\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Cyber-attacks%20on%20hospital%20systems%3A%20a%20narrative%20review&amp;journal=Am.%20J.%20Ger.%20Psychiatry&amp;volume=7&amp;pages=30-39&amp;publication_year=2025&amp;author=Li%2CS&amp;author=Surineni%2CK&amp;author=Prabhakaran%2CN\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR39\">Bowsher, G., Sullivan, R. &amp; Lentzos, F. Tackling health disinformation in conflict settings. Lancet 405, 1052 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/S0140-6736(25)00456-8\" data-track-item_id=\"10.1016\/S0140-6736(25)00456-8\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2FS0140-6736%2825%2900456-8\" aria-label=\"Article reference 39\" data-doi=\"10.1016\/S0140-6736(25)00456-8\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40157794\" aria-label=\"PubMed reference 39\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 39\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Tackling%20health%20disinformation%20in%20conflict%20settings&amp;journal=Lancet&amp;doi=10.1016%2FS0140-6736%2825%2900456-8&amp;volume=405&amp;publication_year=2025&amp;author=Bowsher%2CG&amp;author=Sullivan%2CR&amp;author=Lentzos%2CF\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR40\">Menz, B. D., Modi, N. D., Sorich, M. J. &amp; Hopkins, A. M. Health disinformation use case highlighting the urgent need for artificial intelligence vigilance: weapons of mass disinformation: weapons of mass disinformation. JAMA Intern. Med. 184, 92\u201396 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1001\/jamainternmed.2023.5947\" data-track-item_id=\"10.1001\/jamainternmed.2023.5947\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1001%2Fjamainternmed.2023.5947\" aria-label=\"Article reference 40\" data-doi=\"10.1001\/jamainternmed.2023.5947\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=37955873\" aria-label=\"PubMed reference 40\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 40\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Health%20disinformation%20use%20case%20highlighting%20the%20urgent%20need%20for%20artificial%20intelligence%20vigilance%3A%20weapons%20of%20mass%20disinformation%3A%20weapons%20of%20mass%20disinformation&amp;journal=JAMA%20Intern.%20Med.&amp;doi=10.1001%2Fjamainternmed.2023.5947&amp;volume=184&amp;pages=92-96&amp;publication_year=2024&amp;author=Menz%2CBD&amp;author=Modi%2CND&amp;author=Sorich%2CMJ&amp;author=Hopkins%2CAM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR41\">Perakslis, E. D., Ranney, M. L. &amp; Goldsack, J. C. Characterizing cyber harms from digital health. Nat. Med. 29, 528\u2013531 (2023).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41591-022-02167-6\" data-track-item_id=\"10.1038\/s41591-022-02167-6\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41591-022-02167-6\" aria-label=\"Article reference 41\" data-doi=\"10.1038\/s41591-022-02167-6\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB3sXjtVKgsL0%3D\" aria-label=\"CAS reference 41\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=36759672\" aria-label=\"PubMed reference 41\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 41\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Characterizing%20cyber%20harms%20from%20digital%20health&amp;journal=Nat.%20Med.&amp;doi=10.1038%2Fs41591-022-02167-6&amp;volume=29&amp;pages=528-531&amp;publication_year=2023&amp;author=Perakslis%2CED&amp;author=Ranney%2CML&amp;author=Goldsack%2CJC\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR42\">Nweke, L. O. Using the CIA and AAA Models to Explain Cybersecurity Activities. Preprint at <a href=\"https:\/\/pmworldlibrary.net\/wp-content\/uploads\/2017\/05\/171126-Nweke-Using-CIA-and-AAA-Models-to-explain-Cybersecurity.pdf\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/pmworldlibrary.net\/wp-content\/uploads\/2017\/05\/171126-Nweke-Using-CIA-and-AAA-Models-to-explain-Cybersecurity.pdf\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/pmworldlibrary.net\/wp-content\/uploads\/2017\/05\/171126-Nweke-Using-CIA-and-AAA-Models-to-explain-Cybersecurity.pdf<\/a> (2017).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR43\">Nasr, M. et al. Scalable Extraction of Training Data from Algined, Production Language Models. ICLR 2025 <a href=\"https:\/\/openreview.net\/forum?id=vjel3nWP2a\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/openreview.net\/forum?id=vjel3nWP2a\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/openreview.net\/forum?id=vjel3nWP2a<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR44\">Jha, R., Zhang, C., Shmatikov, V. &amp; Morris, J. X. Harnessing the universal geometry of embeddings. NeurIPS 2025 Conference <a href=\"https:\/\/neurips.cc\/virtual\/2025\/loc\/san-diego\/poster\/116441\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/neurips.cc\/virtual\/2025\/loc\/san-diego\/poster\/116441\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/neurips.cc\/virtual\/2025\/loc\/san-diego\/poster\/116441<\/a> (NeurIPS, 2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR45\">Finlayson, S. G., Chung, H. W., Kohane, I. S. &amp; Beam, A. L. Adversarial attacks against medical deep learning systems. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.1804.05296\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.1804.05296\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.1804.05296<\/a> (2018). This early and influential work identifies adversarial attacks as a fundamental risk for medical machine learning in healthcare.<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR46\">Chakraborty, A., Alam, M., Dey, V., Chattopadhyay, A. &amp; Mukhopadhyay, D. Adversarial attacks and defences: a survey. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.1810.00069\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.1810.00069\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.1810.00069<\/a> (2018).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR47\">Xu, H. et al. Adversarial attacks and defenses in images, graphs and text: a review. Int. J. Autom. Comput. 17, 151\u2013178 (2020).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"noopener nofollow\" data-track-label=\"10.1007\/s11633-019-1211-x\" data-track-item_id=\"10.1007\/s11633-019-1211-x\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/link.springer.com\/doi\/10.1007\/s11633-019-1211-x\" aria-label=\"Article reference 47\" data-doi=\"10.1007\/s11633-019-1211-x\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 47\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Adversarial%20attacks%20and%20defenses%20in%20images%2C%20graphs%20and%20text%3A%20a%20review&amp;journal=Int.%20J.%20Autom.%20Comput.&amp;doi=10.1007%2Fs11633-019-1211-x&amp;volume=17&amp;pages=151-178&amp;publication_year=2020&amp;author=Xu%2CH\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR48\">Barreno, M., Nelson, B., Joseph, A. D. &amp; Tygar, J. D. The security of machine learning. Mach. Learn. 81, 121\u2013148 (2010).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"noopener nofollow\" data-track-label=\"10.1007\/s10994-010-5188-5\" data-track-item_id=\"10.1007\/s10994-010-5188-5\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/link.springer.com\/doi\/10.1007\/s10994-010-5188-5\" aria-label=\"Article reference 48\" data-doi=\"10.1007\/s10994-010-5188-5\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2010MLear..81..121B\" aria-label=\"ADS reference 48\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"mathscinet reference\" data-track-action=\"mathscinet reference\" href=\"http:\/\/www.ams.org\/mathscinet-getitem?mr=3108177\" aria-label=\"MathSciNet reference 48\" target=\"_blank\">MathSciNet<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 48\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=The%20security%20of%20machine%20learning&amp;journal=Mach.%20Learn.&amp;doi=10.1007%2Fs10994-010-5188-5&amp;volume=81&amp;pages=121-148&amp;publication_year=2010&amp;author=Barreno%2CM&amp;author=Nelson%2CB&amp;author=Joseph%2CAD&amp;author=Tygar%2CJD\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR49\">OWASP Top 10 for LLM Applications 2025 <a href=\"https:\/\/genai.owasp.org\/resource\/owasp-top-10-for-llm-applications-2025\/\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/genai.owasp.org\/resource\/owasp-top-10-for-llm-applications-2025\/\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/genai.owasp.org\/resource\/owasp-top-10-for-llm-applications-2025\/<\/a> (OWASP, 2024). This resource provides a continuously updated overview of the most critical safety and security risks in LLMs.<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR50\">CVSS v4.0 Specification Document. FIRST <a href=\"https:\/\/www.first.org\/cvss\/specification-document\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/www.first.org\/cvss\/specification-document\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.first.org\/cvss\/specification-document<\/a> (2023).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR51\">Kaviani, S., Han, K. J. &amp; Sohn, I. Adversarial attacks and defenses on AI in medical imaging informatics: A survey. Expert Syst. Appl. 198, 116815 (2022).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/j.eswa.2022.116815\" data-track-item_id=\"10.1016\/j.eswa.2022.116815\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2Fj.eswa.2022.116815\" aria-label=\"Article reference 51\" data-doi=\"10.1016\/j.eswa.2022.116815\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 51\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Adversarial%20attacks%20and%20defenses%20on%20AI%20in%20medical%20imaging%20informatics%3A%20A%20survey&amp;journal=Expert%20Syst.%20Appl.&amp;doi=10.1016%2Fj.eswa.2022.116815&amp;volume=198&amp;publication_year=2022&amp;author=Kaviani%2CS&amp;author=Han%2CKJ&amp;author=Sohn%2CI\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR52\">Nagaraja, N. &amp; Bahsi, H. Cyber threat modeling of an LLM-based healthcare system. In Proc. 11th International Conference on Information Systems Security and Privacy 325\u2013336 (SCITEPRESS, 2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR53\">Li, M. Q. &amp; Fung, B. C. M. Security concerns for large language models: a survey. Journal of Information Security and Applications, Volume 95, 2025 <a href=\"https:\/\/doi.org\/10.1016\/j.jisa.2025.104284\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.1016\/j.jisa.2025.104284\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.1016\/j.jisa.2025.104284<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR54\">Reason, J. Human error: models and management. Brit. Med. J. 320, 768\u2013770 (2000).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1136\/bmj.320.7237.768\" data-track-item_id=\"10.1136\/bmj.320.7237.768\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1136%2Fbmj.320.7237.768\" aria-label=\"Article reference 54\" data-doi=\"10.1136\/bmj.320.7237.768\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:STN:280:DC%2BD3c7osFChtQ%3D%3D\" aria-label=\"CAS reference 54\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=10720363\" aria-label=\"PubMed reference 54\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC1117770\" aria-label=\"PubMed Central reference 54\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 54\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Human%20error%3A%20models%20and%20management&amp;journal=Brit.%20Med.%20J.&amp;doi=10.1136%2Fbmj.320.7237.768&amp;volume=320&amp;pages=768-770&amp;publication_year=2000&amp;author=Reason%2CJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR55\">Joint Task Force Transformation Initiative. Guide for Conducting Risk Assessments <a href=\"https:\/\/doi.org\/10.6028\/nist.sp.800-30r1\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.6028\/nist.sp.800-30r1\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.6028\/nist.sp.800-30r1<\/a> (US Department of Commerce, 2012).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR56\">Ashkenazy, S. Why time-to-market is overtaking device security as a top priority. Cybellum <a href=\"https:\/\/cybellum.com\/blog\/why-time-to-market-is-overtaking-device-security-as-a-top-priority-in-2025\/\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/cybellum.com\/blog\/why-time-to-market-is-overtaking-device-security-as-a-top-priority-in-2025\/\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/cybellum.com\/blog\/why-time-to-market-is-overtaking-device-security-as-a-top-priority-in-2025\/<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR57\">Piao, Y., Li, J. &amp; Woods, D. W. Measuring the vulnerability disclosure policies of AI vendors. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2509.06136\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2509.06136\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2509.06136<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR58\">Bommasani, R. et al. The 2024 Foundation Model Transparency Index. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2407.12929\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2407.12929\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2407.12929<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR59\">Anthropic\u2019s responsible scaling policy. Anthropic <a href=\"https:\/\/www.anthropic.com\/rsp-updates\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/www.anthropic.com\/rsp-updates\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.anthropic.com\/rsp-updates<\/a> (2026).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR60\">Lindsey, J. et al. On the biology of a large language model. Anthropic <a href=\"https:\/\/transformer-circuits.pub\/2025\/attribution-graphs\/biology.html\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/transformer-circuits.pub\/2025\/attribution-graphs\/biology.html\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/transformer-circuits.pub\/2025\/attribution-graphs\/biology.html<\/a> (2025). This work provides an extensive analysis of internal mechanisms in LLMs and introduces open-source tools to advance mechanistic interpretability.<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR61\">Petri: an open-source auditing tool to accelerate AI safety research. Anthropic <a href=\"https:\/\/www.anthropic.com\/research\/petri-open-source-auditing\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/www.anthropic.com\/research\/petri-open-source-auditing\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.anthropic.com\/research\/petri-open-source-auditing<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR62\">Joglekar, M. et al. Training LLMs for honesty via confessions. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2512.08093\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2512.08093\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2512.08093<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR63\">Villalobos, P. et al. Position: will we run out of data? Limits of LLM scaling based on human-generated data. ICML 235, 49523\u201349544 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 63\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Position%3A%20will%20we%20run%20out%20of%20data%3F%20Limits%20of%20LLM%20scaling%20based%20on%20human-generated%20data&amp;journal=ICML&amp;volume=235&amp;pages=49523-49544&amp;publication_year=2024&amp;author=Villalobos%2CP\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR64\">Alber, D. A. et al. Medical large language models are vulnerable to data-poisoning attacks. Nat. Med. 31, 618\u2013626 (2025). This study demonstrates that even minimal data poisoning can induce clinically harmful behaviour in medical LLMs.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41591-024-03445-1\" data-track-item_id=\"10.1038\/s41591-024-03445-1\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41591-024-03445-1\" aria-label=\"Article reference 64\" data-doi=\"10.1038\/s41591-024-03445-1\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXhtVWqsr8%3D\" aria-label=\"CAS reference 64\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39779928\" aria-label=\"PubMed reference 64\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11835729\" aria-label=\"PubMed Central reference 64\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 64\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Medical%20large%20language%20models%20are%20vulnerable%20to%20data-poisoning%20attacks&amp;journal=Nat.%20Med.&amp;doi=10.1038%2Fs41591-024-03445-1&amp;volume=31&amp;pages=618-626&amp;publication_year=2025&amp;author=Alber%2CDA\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR65\">Hubinger, E. et al. Sleeper agents: training deceptive LLMs that persist through safety training. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2401.05566\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2401.05566\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2401.05566<\/a> (2024). This study characterizes difficult-to-detect and persistent \u2018sleeper agent\u2019 backdoor behaviours in LLMs.<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR66\">Souri, H., Fowl, L. H., Chellappa, R., Goldblum, M. &amp; Goldstein, T. in Advances in Neural Information Processing Systems 35 (eds Koyejo, S. et al.) 19165\u201319178 (NeurIPS, 2022).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR67\">Carlini, N. et al. Poisoning Web-Scale Training Datasets is Practical. in 2024 IEEE Symposium on Security and Privacy Vol. 29, 407\u2013425 (IEEE, 2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR68\">Clusmann, J. et al. Incidental prompt injections on vision\u2013language models in real-life histopathology. NEJM AI 2, AIcs2500078 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1056\/AIcs2500078\" data-track-item_id=\"10.1056\/AIcs2500078\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1056%2FAIcs2500078\" aria-label=\"Article reference 68\" data-doi=\"10.1056\/AIcs2500078\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 68\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Incidental%20prompt%20injections%20on%20vision%E2%80%93language%20models%20in%20real-life%20histopathology&amp;journal=NEJM%20AI&amp;doi=10.1056%2FAIcs2500078&amp;volume=2&amp;publication_year=2025&amp;author=Clusmann%2CJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR69\">Han, T. et al. Medical large language models are susceptible to targeted misinformation attacks. npj Digit. Med. 7, 288 (2024). This study demonstrates the feasibility of targeted model weight manipulation to introduce incorrect biomedical information into LLMs.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41746-024-01282-7\" data-track-item_id=\"10.1038\/s41746-024-01282-7\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41746-024-01282-7\" aria-label=\"Article reference 69\" data-doi=\"10.1038\/s41746-024-01282-7\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39443664\" aria-label=\"PubMed reference 69\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11499642\" aria-label=\"PubMed Central reference 69\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 69\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Medical%20large%20language%20models%20are%20susceptible%20to%20targeted%20misinformation%20attacks&amp;journal=npj%20Digit.%20Med.&amp;doi=10.1038%2Fs41746-024-01282-7&amp;volume=7&amp;publication_year=2024&amp;author=Han%2CT\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR70\">Rao, P. S. B., \u0160\u0107epanovi\u0107, S., Jayagopi, D. B., Cherubini, M. &amp; Quercia, D. The AI model risk catalog: what developers and researchers miss about real-world AI harms. In Proc. AAAI\/ACM Conference on AI, Ethics, and Society Vol. 8, 2163\u20132150 (AAAI\/ACM, 2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR71\">Carlini, N. et al. Poisoning web-scale training datasets is practical. ar5iv <a href=\"https:\/\/ar5iv.labs.arxiv.org\/html\/2302.10149\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/ar5iv.labs.arxiv.org\/html\/2302.10149\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/ar5iv.labs.arxiv.org\/html\/2302.10149<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR72\">Clusmann, J. et al. Prompt injection attacks on vision language models in oncology. Nat. Commun. 16, 1239 (2025). This study demonstrates prompt injection attacks on LLMs through hidden prompt instructions on medical imaging data.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41467-024-55631-x\" data-track-item_id=\"10.1038\/s41467-024-55631-x\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41467-024-55631-x\" aria-label=\"Article reference 72\" data-doi=\"10.1038\/s41467-024-55631-x\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2025NatCo..16.1239C\" aria-label=\"ADS reference 72\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXjtVaktb0%3D\" aria-label=\"CAS reference 72\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39890777\" aria-label=\"PubMed reference 72\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11785991\" aria-label=\"PubMed Central reference 72\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 72\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Prompt%20injection%20attacks%20on%20vision%20language%20models%20in%20oncology&amp;journal=Nat.%20Commun.&amp;doi=10.1038%2Fs41467-024-55631-x&amp;volume=16&amp;publication_year=2025&amp;author=Clusmann%2CJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR73\">Zhang, Z., Qadir, M.I., Carstens, M. et al. Prompt injection attacks on vision-language models for surgical decision support. npj Digit. Surg. 1, 15 <a href=\"https:\/\/doi.org\/10.1038\/s44484-026-00014-6\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.1038\/s44484-026-00014-6\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.1038\/s44484-026-00014-6<\/a> (2026).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR74\">Lapuschkin, S. et al. Unmasking Clever Hans predictors and assessing what machines really learn. Nat. Commun. 10, 1096 (2019).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41467-019-08987-4\" data-track-item_id=\"10.1038\/s41467-019-08987-4\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41467-019-08987-4\" aria-label=\"Article reference 74\" data-doi=\"10.1038\/s41467-019-08987-4\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2019NatCo..10.1096L\" aria-label=\"ADS reference 74\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=30858366\" aria-label=\"PubMed reference 74\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC6411769\" aria-label=\"PubMed Central reference 74\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 74\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Unmasking%20Clever%20Hans%20predictors%20and%20assessing%20what%20machines%20really%20learn&amp;journal=Nat.%20Commun.&amp;doi=10.1038%2Fs41467-019-08987-4&amp;volume=10&amp;publication_year=2019&amp;author=Lapuschkin%2CS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR75\">Hou, G. et al. Evaluating robustness of large audio language models to audio injection: an empirical study. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2505.19598\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2505.19598\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2505.19598<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR76\">Rajeev, M. et al. Cats confuse reasoning LLM: query agnostic adversarial triggers for reasoning models. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2503.01781\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2503.01781\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2503.01781<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR77\">Alizadeh, M., Samei, Z., Stetsenko, D. &amp; Gilardi, F. Simple prompt injection attacks can leak personal data observed by LLM agents during task execution. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2506.01055\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2506.01055\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2506.01055<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR78\">Meincke, L. et al. Persuading large language models to comply with objectionable requests. Proc. Natl Acad. Sci. USA 123, e2535868123 (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR79\">Zeng, Y. et al. How Johnny can persuade LLMs to jailbreak them: rethinking persuasion to challenge AI safety by humanizing LLMs. In Proc. 62nd Meeting of the Association for Computational Linguistics Vol. 1, 14322\u201314350 (ACL, 2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR80\">Schoene, A. M. &amp; Canca, C. \u2018For argument\u2019s sake, show me how to harm myself!\u2019: Jailbreaking LLMs in suicide and self-harm contexts. In 2025 IEEE International Symposium on Technology and Society <a href=\"https:\/\/doi.org\/10.1109\/ISTAS65609.2025.11269647\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.1109\/ISTAS65609.2025.11269647\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.1109\/ISTAS65609.2025.11269647<\/a> (IEEE, 2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR81\">Mehrotra, A. et al. Tree of Attacks: Jailbreaking Black-Box LLMs Automatically. In Advances in Neural Information Processing Systems 37 (eds Globerson, A. et al.) 61065\u201361105 (NeurIPS, 2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR82\">Jiang, F. et al. ArtPrompt: ASCII art-based jailbreak attacks against aligned LLMs. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2402.11753\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2402.11753\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2402.11753<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR83\">Zhang, H., Lou, Q. &amp; Wang, Y. Towards safe AI clinicians: a comprehensive study on large language model jailbreaking in healthcare. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2501.18632\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2501.18632\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2501.18632<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR84\">Sharma, M. et al. Constitutional classifiers: defending against universal jailbreaks across thousands of hours of red teaming. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2501.18837\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2501.18837\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2501.18837<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR85\">Yuan, Y. et al. From hard refusals to safe-completions: toward output-centric safety training. Superintelligence <a href=\"https:\/\/doi.org\/10.70777\/si.v2i6.15625\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.70777\/si.v2i6.15625\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.70777\/si.v2i6.15625<\/a> (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.70777\/si.v2i6.15625\" data-track-item_id=\"10.70777\/si.v2i6.15625\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.70777%2Fsi.v2i6.15625\" aria-label=\"Article reference 85\" data-doi=\"10.70777\/si.v2i6.15625\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 85\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=From%20hard%20refusals%20to%20safe-completions%3A%20toward%20output-centric%20safety%20training&amp;journal=Superintelligence&amp;doi=10.70777%2Fsi.v2i6.15625&amp;publication_year=2025&amp;author=Yuan%2CY\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR86\">Zhang, Y., Juels, A., Reiter, M. K. &amp; Ristenpart, T. Cross-VM side channels and their use to extract private keys. In Proc. 2012 ACM Conference on Computer and Communications Security <a href=\"https:\/\/doi.org\/10.1145\/2382196.2382230\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.1145\/2382196.2382230\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.1145\/2382196.2382230<\/a> (ACM, 2012).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR87\">Karim, H., Gupta, D. &amp; Sitharaman, S. Securing LLM workloads with NIST AI RMF in the internet of robotic things. IEEE Access 13, 69631\u201369649 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1109\/ACCESS.2025.3561235\" data-track-item_id=\"10.1109\/ACCESS.2025.3561235\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1109%2FACCESS.2025.3561235\" aria-label=\"Article reference 87\" data-doi=\"10.1109\/ACCESS.2025.3561235\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 87\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Securing%20LLM%20workloads%20with%20NIST%20AI%20RMF%20in%20the%20internet%20of%20robotic%20things&amp;journal=IEEE%20Access&amp;doi=10.1109%2FACCESS.2025.3561235&amp;volume=13&amp;pages=69631-69649&amp;publication_year=2025&amp;author=Karim%2CH&amp;author=Gupta%2CD&amp;author=Sitharaman%2CS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR88\">Huang, H., Meng, T. &amp; Jia, W. Joint optimization of prompt security and system performance in Edge-Cloud LLM systems. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2501.18663\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2501.18663\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2501.18663<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR89\">Schmotz, D., Abdelnabi, S. &amp; Andriushchenko, M. Agent skills enable a new class of realistic and trivially simple prompt injections. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2510.26328\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2510.26328\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2510.26328<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR90\">Narajala, V. S. &amp; Habler, I. Enterprise-grade security for the model context protocol (MCP): frameworks and mitigation strategies. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2504.08623\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2504.08623\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2504.08623<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR91\">Thoonsen, A. C. et al. Stimulating implementation of clinical practice guidelines in hospital care from a central guideline organization perspective: A systematic review. Health Policy 148, 105135 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/j.healthpol.2024.105135\" data-track-item_id=\"10.1016\/j.healthpol.2024.105135\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2Fj.healthpol.2024.105135\" aria-label=\"Article reference 91\" data-doi=\"10.1016\/j.healthpol.2024.105135\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39128438\" aria-label=\"PubMed reference 91\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 91\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Stimulating%20implementation%20of%20clinical%20practice%20guidelines%20in%20hospital%20care%20from%20a%20central%20guideline%20organization%20perspective%3A%20A%20systematic%20review&amp;journal=Health%20Policy&amp;doi=10.1016%2Fj.healthpol.2024.105135&amp;volume=148&amp;publication_year=2024&amp;author=Thoonsen%2CAC\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR92\">Clusmann, J. et al. The barriers for uptake of artificial intelligence in hepatology and how to overcome them. J. Hepatol. 83, 1410\u20131426 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/j.jhep.2025.07.003\" data-track-item_id=\"10.1016\/j.jhep.2025.07.003\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2Fj.jhep.2025.07.003\" aria-label=\"Article reference 92\" data-doi=\"10.1016\/j.jhep.2025.07.003\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40920593\" aria-label=\"PubMed reference 92\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 92\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=The%20barriers%20for%20uptake%20of%20artificial%20intelligence%20in%20hepatology%20and%20how%20to%20overcome%20them&amp;journal=J.%20Hepatol.&amp;doi=10.1016%2Fj.jhep.2025.07.003&amp;volume=83&amp;pages=1410-1426&amp;publication_year=2025&amp;author=Clusmann%2CJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR93\">Wong, E. Y. T. et al. ESMO guidance on the use of large language models in clinical practice (ELCAP). Ann. Oncol. 36, 1447\u20131457 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/j.annonc.2025.09.001\" data-track-item_id=\"10.1016\/j.annonc.2025.09.001\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2Fj.annonc.2025.09.001\" aria-label=\"Article reference 93\" data-doi=\"10.1016\/j.annonc.2025.09.001\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:STN:280:DC%2BA3M7jvFSmsg%3D%3D\" aria-label=\"CAS reference 93\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=41111032\" aria-label=\"PubMed reference 93\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 93\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=ESMO%20guidance%20on%20the%20use%20of%20large%20language%20models%20in%20clinical%20practice%20%28ELCAP%29&amp;journal=Ann.%20Oncol.&amp;doi=10.1016%2Fj.annonc.2025.09.001&amp;volume=36&amp;pages=1447-1457&amp;publication_year=2025&amp;author=Wong%2CEYT\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR94\">Sarthi, P. et al. RAPTOR: recursive abstractive processing for tree-organized retrieval. ICLR <a href=\"https:\/\/iclr.cc\/virtual\/2024\/poster\/19034\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/iclr.cc\/virtual\/2024\/poster\/19034\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/iclr.cc\/virtual\/2024\/poster\/19034<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR95\">Greshake, K. et al. Not what you\u2019ve signed up for: compromising real-world LLM-integrated applications with indirect prompt injection. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2302.12173\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2302.12173\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2302.12173<\/a> (2023).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR96\">Continella, A., Polino, M., Pogliani, M. &amp; Zanero, S. There\u2019s a hole in that bucket!: a large-scale analysis of misconfigured S3 buckets. In Proc. 34th Annual Computer Security Applications Conference 702\u2013711 (ACM, 2018).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR97\">Cruz, F. &amp; Lombrozo, T. How laypeople evaluate scientific explanations containing jargon. Nat. Hum. Behav. 9, 2038\u20132053 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41562-025-02227-0\" data-track-item_id=\"10.1038\/s41562-025-02227-0\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41562-025-02227-0\" aria-label=\"Article reference 97\" data-doi=\"10.1038\/s41562-025-02227-0\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40506550\" aria-label=\"PubMed reference 97\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 97\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=How%20laypeople%20evaluate%20scientific%20explanations%20containing%20jargon&amp;journal=Nat.%20Hum.%20Behav.&amp;doi=10.1038%2Fs41562-025-02227-0&amp;volume=9&amp;pages=2038-2053&amp;publication_year=2025&amp;author=Cruz%2CF&amp;author=Lombrozo%2CT\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR98\">Measuring the Persuasiveness of Language Models. Anthropic <a href=\"https:\/\/www.anthropic.com\/news\/measuring-model-persuasiveness\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/www.anthropic.com\/news\/measuring-model-persuasiveness\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.anthropic.com\/news\/measuring-model-persuasiveness<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR99\">Salvi, F., Horta Ribeiro, M., Gallotti, R. &amp; West, R. On the conversational persuasiveness of GPT-4. Nat. Hum. Behav. 9, 1645\u20131653 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41562-025-02194-6\" data-track-item_id=\"10.1038\/s41562-025-02194-6\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41562-025-02194-6\" aria-label=\"Article reference 99\" data-doi=\"10.1038\/s41562-025-02194-6\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40389594\" aria-label=\"PubMed reference 99\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12367540\" aria-label=\"PubMed Central reference 99\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 99\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=On%20the%20conversational%20persuasiveness%20of%20GPT-4&amp;journal=Nat.%20Hum.%20Behav.&amp;doi=10.1038%2Fs41562-025-02194-6&amp;volume=9&amp;pages=1645-1653&amp;publication_year=2025&amp;author=Salvi%2CF&amp;author=Horta%20Ribeiro%2CM&amp;author=Gallotti%2CR&amp;author=West%2CR\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR100\">Shekar, S., Pataranutaporn, P., Sarabu, C., Cecchi, G. A. &amp; Maes, P. People overtrust AI-generated medical advice despite low accuracy. NEJM AI 2, AIoa2300015 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1056\/AIoa2300015\" data-track-item_id=\"10.1056\/AIoa2300015\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1056%2FAIoa2300015\" aria-label=\"Article reference 100\" data-doi=\"10.1056\/AIoa2300015\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 100\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=People%20overtrust%20AI-generated%20medical%20advice%20despite%20low%20accuracy&amp;journal=NEJM%20AI&amp;doi=10.1056%2FAIoa2300015&amp;volume=2&amp;publication_year=2025&amp;author=Shekar%2CS&amp;author=Pataranutaporn%2CP&amp;author=Sarabu%2CC&amp;author=Cecchi%2CGA&amp;author=Maes%2CP\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR101\">Li, W. et al. Can a large language model be a gaslighter? In The 13th International Conference on Learning Representations (eds Yue, Y. et al.) <a href=\"https:\/\/proceedings.iclr.cc\/paper_files\/paper\/2025\/file\/0769598fdeb4f23ee86fec1bc0777f44-Paper-Conference.pdf\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/proceedings.iclr.cc\/paper_files\/paper\/2025\/file\/0769598fdeb4f23ee86fec1bc0777f44-Paper-Conference.pdf\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/proceedings.iclr.cc\/paper_files\/paper\/2025\/file\/0769598fdeb4f23ee86fec1bc0777f44-Paper-Conference.pdf<\/a> (ICLR, 2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR102\">Yeung, J. A., Dalmasso, J., Foschini, L., Dobson, R. J. B. &amp; Kraljevic, Z. The psychogenic machine: simulating AI psychosis, delusion reinforcement and harm enablement in large language models. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2509.109702025\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2509.109702025\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2509.109702025<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR103\">Turpin, M., Michael, J., Perez, E. &amp; Bowman, S. R. Language models don\u2019t always say what they think: unfaithful explanations in chain-of-thought prompting. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2305.04388\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2305.04388\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2305.04388<\/a> (2023).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR104\">Baker, B. et al. Monitoring reasoning models for misbehavior and the risks of promoting obfuscation. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2503.11926\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2503.11926\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2503.11926<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR105\">Perez, E. et al. Discovering language model behaviors with model-written evaluations. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2212.09251\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2212.09251\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2212.09251<\/a> (2022).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR106\">Xiao, J. et al. On the algorithmic bias of aligning large language models with RLHF: Preference collapse and matching regularization. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2405.16455\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2405.16455\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2405.16455<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR107\">Chen, S. et al. When helpfulness backfires: LLMs and the risk of false medical information due to sycophantic behavior. NPJ Digit. Med. 8, 605 (2025). This study identifies sycophantic behaviour in LLMs as a potential source of false medical information in healthcare settings.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41746-025-02008-z\" data-track-item_id=\"10.1038\/s41746-025-02008-z\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41746-025-02008-z\" aria-label=\"Article reference 107\" data-doi=\"10.1038\/s41746-025-02008-z\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=41107408\" aria-label=\"PubMed reference 107\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12534679\" aria-label=\"PubMed Central reference 107\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 107\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=When%20helpfulness%20backfires%3A%20LLMs%20and%20the%20risk%20of%20false%20medical%20information%20due%20to%20sycophantic%20behavior&amp;journal=NPJ%20Digit.%20Med.&amp;doi=10.1038%2Fs41746-025-02008-z&amp;volume=8&amp;publication_year=2025&amp;author=Chen%2CS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR108\">Kalai, A. T., Nachum, O., Vempala, S. S. &amp; Zhang, E. Why language models hallucinate. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2509.04664\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2509.04664\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2509.04664<\/a> (2025). This study characterizes hallucinations in LLMs as a consequence of misaligned training objectives that reward plausible guessing over admitting uncertainty.<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR109\">Sharma, M. et al. Towards understanding sycophancy in language models. In The 12th International Conference on Learning Representations (eds Kim, B. et al.) <a href=\"http:\/\/proceedings.iclr.cc\/paper_files\/paper\/2024\/file\/0105f7972202c1d4fb817da9f21a9663-Paper-Conference.pdf\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"http:\/\/proceedings.iclr.cc\/paper_files\/paper\/2024\/file\/0105f7972202c1d4fb817da9f21a9663-Paper-Conference.pdf\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/proceedings.iclr.cc\/paper_files\/paper\/2024\/file\/0105f7972202c1d4fb817da9f21a9663-Paper-Conference.pdf<\/a> (ICLR, 2023).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR110\">Cheng, M. et al. Sycophantic AI decreases prosocial intentions and promotes dependence. Science 391, eaec8352 (2026).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1126\/science.aec8352\" data-track-item_id=\"10.1126\/science.aec8352\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1126%2Fscience.aec8352\" aria-label=\"Article reference 110\" data-doi=\"10.1126\/science.aec8352\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB28XosVKntLY%3D\" aria-label=\"CAS reference 110\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=41886588\" aria-label=\"PubMed reference 110\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 110\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Sycophantic%20AI%20decreases%20prosocial%20intentions%20and%20promotes%20dependence&amp;journal=Science&amp;doi=10.1126%2Fscience.aec8352&amp;volume=391&amp;publication_year=2026&amp;author=Cheng%2CM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR111\">Miton, H., Claidi\u00e8re, N. &amp; Mercier, H. Universal cognitive mechanisms explain the cultural success of bloodletting. Evol. Hum. Behav. 36, 303\u2013312 (2015).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/j.evolhumbehav.2015.01.003\" data-track-item_id=\"10.1016\/j.evolhumbehav.2015.01.003\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2Fj.evolhumbehav.2015.01.003\" aria-label=\"Article reference 111\" data-doi=\"10.1016\/j.evolhumbehav.2015.01.003\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 111\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Universal%20cognitive%20mechanisms%20explain%20the%20cultural%20success%20of%20bloodletting&amp;journal=Evol.%20Hum.%20Behav.&amp;doi=10.1016%2Fj.evolhumbehav.2015.01.003&amp;volume=36&amp;pages=303-312&amp;publication_year=2015&amp;author=Miton%2CH&amp;author=Claidi%C3%A8re%2CN&amp;author=Mercier%2CH\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR112\">Saposnik, G., Redelmeier, D., Ruff, C. C. &amp; Tobler, P. N. Cognitive biases associated with medical decisions: a systematic review. BMC Med. Inform. Decis. Mak. 16, 138 (2016).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"noopener nofollow\" data-track-label=\"10.1186\/s12911-016-0377-1\" data-track-item_id=\"10.1186\/s12911-016-0377-1\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/link.springer.com\/doi\/10.1186\/s12911-016-0377-1\" aria-label=\"Article reference 112\" data-doi=\"10.1186\/s12911-016-0377-1\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=27809908\" aria-label=\"PubMed reference 112\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC5093937\" aria-label=\"PubMed Central reference 112\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 112\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Cognitive%20biases%20associated%20with%20medical%20decisions%3A%20a%20systematic%20review&amp;journal=BMC%20Med.%20Inform.%20Decis.%20Mak.&amp;doi=10.1186%2Fs12911-016-0377-1&amp;volume=16&amp;publication_year=2016&amp;author=Saposnik%2CG&amp;author=Redelmeier%2CD&amp;author=Ruff%2CCC&amp;author=Tobler%2CPN\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR113\">Huang, L. et al. A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Trans. Inf. Syst. 43, 42 (2025). This study provides a systematic analysis and taxonomy of hallucinations in LLMs.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1145\/3703155\" data-track-item_id=\"10.1145\/3703155\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1145%2F3703155\" aria-label=\"Article reference 113\" data-doi=\"10.1145\/3703155\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 113\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=A%20survey%20on%20hallucination%20in%20large%20language%20models%3A%20principles%2C%20taxonomy%2C%20challenges%2C%20and%20open%20questions&amp;journal=ACM%20Trans.%20Inf.%20Syst.&amp;doi=10.1145%2F3703155&amp;volume=43&amp;publication_year=2025&amp;author=Huang%2CL\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR114\">Ferber, D. et al. In-context learning enables multimodal large language models to classify cancer pathology images. Nat. Commun. 15, 10104 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41467-024-51465-9\" data-track-item_id=\"10.1038\/s41467-024-51465-9\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41467-024-51465-9\" aria-label=\"Article reference 114\" data-doi=\"10.1038\/s41467-024-51465-9\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2024NatCo..1510104F\" aria-label=\"ADS reference 114\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2cXisF2hurnE\" aria-label=\"CAS reference 114\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39572531\" aria-label=\"PubMed reference 114\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11582649\" aria-label=\"PubMed Central reference 114\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 114\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=In-context%20learning%20enables%20multimodal%20large%20language%20models%20to%20classify%20cancer%20pathology%20images&amp;journal=Nat.%20Commun.&amp;doi=10.1038%2Fs41467-024-51465-9&amp;volume=15&amp;publication_year=2024&amp;author=Ferber%2CD\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR115\">Gourabathina, A., Gerych, W., Pan, E. &amp; Ghassemi, M. The medium is the message: how non-clinical information shapes clinical decisions in LLMs. In Proc. 2025 ACM Conference on Fairness, Accountability, and Transparency 1805\u20131828 (ACM, 2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR116\">Coda-Forno, J. et al. Inducing anxiety in large language models can induce bias. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2304.11111\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2304.11111\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2304.11111<\/a> (2023).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR117\">Corbeil, J.-P., Kim, M., Sordoni, A., Beaulieu, F. &amp; Vozila, P. Medical red teaming protocol of language models: on the importance of user perspectives in healthcare settings. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2507.07248\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2507.07248\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2507.07248<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR118\">Callahan, A. et al. Standing on FURM ground: a framework for evaluating fair, useful, and reliable AI models in health care systems. NEJM Catal. Innov. Care Deliv. <a href=\"https:\/\/doi.org\/10.1056\/CAT.24.0131\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.1056\/CAT.24.0131\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.1056\/CAT.24.0131<\/a> (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1056\/CAT.24.0131\" data-track-item_id=\"10.1056\/CAT.24.0131\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1056%2FCAT.24.0131\" aria-label=\"Article reference 118\" data-doi=\"10.1056\/CAT.24.0131\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 118\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Standing%20on%20FURM%20ground%3A%20a%20framework%20for%20evaluating%20fair%2C%20useful%2C%20and%20reliable%20AI%20models%20in%20health%20care%20systems&amp;journal=NEJM%20Catal.%20Innov.%20Care%20Deliv.&amp;doi=10.1056%2FCAT.24.0131&amp;publication_year=2024&amp;author=Callahan%2CA\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR119\">Huang, Y. et al. Position: TrustLLM: trustworthiness in large language models. In Proc. 41st International Conference on Machine Learning 20166\u201320270 (PMLR, 2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR120\">Gabriel, I. Artificial intelligence, values, and alignment. Minds Mach. 30, 411\u2013437 (2020).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"noopener nofollow\" data-track-label=\"10.1007\/s11023-020-09539-2\" data-track-item_id=\"10.1007\/s11023-020-09539-2\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/link.springer.com\/doi\/10.1007\/s11023-020-09539-2\" aria-label=\"Article reference 120\" data-doi=\"10.1007\/s11023-020-09539-2\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 120\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Artificial%20intelligence%2C%20values%2C%20and%20alignment&amp;journal=Minds%20Mach.&amp;doi=10.1007%2Fs11023-020-09539-2&amp;volume=30&amp;pages=411-437&amp;publication_year=2020&amp;author=Gabriel%2CI\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR121\">M\u00f6kander, J., Schuett, J., Kirk, H. R. &amp; Floridi, L. Auditing large language models: a three-layered approach. AI Ethics 4, 1085\u20131115 (2023).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"noopener nofollow\" data-track-label=\"10.1007\/s43681-023-00289-2\" data-track-item_id=\"10.1007\/s43681-023-00289-2\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/link.springer.com\/doi\/10.1007\/s43681-023-00289-2\" aria-label=\"Article reference 121\" data-doi=\"10.1007\/s43681-023-00289-2\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 121\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Auditing%20large%20language%20models%3A%20a%20three-layered%20approach&amp;journal=AI%20Ethics&amp;doi=10.1007%2Fs43681-023-00289-2&amp;volume=4&amp;pages=1085-1115&amp;publication_year=2023&amp;author=M%C3%B6kander%2CJ&amp;author=Schuett%2CJ&amp;author=Kirk%2CHR&amp;author=Floridi%2CL\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR122\">Wu, D. et al. First, do NOHARM: towards clinically safe large language models. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2512.01241\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2512.01241\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2512.01241<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR123\">Bedi, S. et al. Holistic evaluation of large language models for medical tasks with MedHELM. Nat. Med. 32, 943\u2013951 (2026). This work proposes MedHELM, a comprehensive holistic evaluation suite for LLMs for medicine.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41591-025-04151-2\" data-track-item_id=\"10.1038\/s41591-025-04151-2\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41591-025-04151-2\" aria-label=\"Article reference 123\" data-doi=\"10.1038\/s41591-025-04151-2\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB28XitFKgsLw%3D\" aria-label=\"CAS reference 123\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=41559415\" aria-label=\"PubMed reference 123\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC13267972\" aria-label=\"PubMed Central reference 123\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 123\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Holistic%20evaluation%20of%20large%20language%20models%20for%20medical%20tasks%20with%20MedHELM&amp;journal=Nat.%20Med.&amp;doi=10.1038%2Fs41591-025-04151-2&amp;volume=32&amp;pages=943-951&amp;publication_year=2026&amp;author=Bedi%2CS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR124\">Han, T., Kumar, A., Agarwal, C. &amp; Lakkaraju, H. in ICML 2024 Workshop on Models of Human Feedback for AI Alignment <a href=\"https:\/\/icml.cc\/virtual\/2024\/39424\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/icml.cc\/virtual\/2024\/39424\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/icml.cc\/virtual\/2024\/39424<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR125\">Shihab, I. F., Akter, S. &amp; Sharma, A. Detecting and mitigating reward hacking in Reinforcement Learning systems: A comprehensive empirical study. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2507.05619\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2507.05619\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2507.05619<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR126\">Wu, S. et al. A comparative study on reasoning patterns of OpenAI\u2019s o1 model. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2410.13639\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2410.13639\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2410.13639<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR127\">El, B. &amp; Zou, J. Moloch\u2019s bargain: emergent misalignment when LLMs compete for audiences. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2510.06105\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2510.06105\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2510.06105<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR128\">Betley, J. et al. Training large language models on narrow tasks can lead to broad misalignment. Nature 649, 584\u2013589 (2026). This study characterizes \u2018emergent misalignment\u2019 and shows that that fine-tuning LLMs on narrow tasks can induce broad, unintended and harmful behaviours.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41586-025-09937-5\" data-track-item_id=\"10.1038\/s41586-025-09937-5\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41586-025-09937-5\" aria-label=\"Article reference 128\" data-doi=\"10.1038\/s41586-025-09937-5\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2026Natur.649..584B\" aria-label=\"ADS reference 128\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB28XhslKiu7o%3D\" aria-label=\"CAS reference 128\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=41535488\" aria-label=\"PubMed reference 128\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12804084\" aria-label=\"PubMed Central reference 128\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 128\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Training%20large%20language%20models%20on%20narrow%20tasks%20can%20lead%20to%20broad%20misalignment&amp;journal=Nature&amp;doi=10.1038%2Fs41586-025-09937-5&amp;volume=649&amp;pages=584-589&amp;publication_year=2026&amp;author=Betley%2CJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR129\">Truhn, D., Reis-Filho, J. S. &amp; Kather, J. N. Large language models should be used as scientific reasoning engines, not knowledge databases. Nat. Med. 29, 2983\u20132984 (2023).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41591-023-02594-z\" data-track-item_id=\"10.1038\/s41591-023-02594-z\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41591-023-02594-z\" aria-label=\"Article reference 129\" data-doi=\"10.1038\/s41591-023-02594-z\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB3sXitFGqt7fN\" aria-label=\"CAS reference 129\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=37853138\" aria-label=\"PubMed reference 129\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 129\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Large%20language%20models%20should%20be%20used%20as%20scientific%20reasoning%20engines%2C%20not%20knowledge%20databases&amp;journal=Nat.%20Med.&amp;doi=10.1038%2Fs41591-023-02594-z&amp;volume=29&amp;pages=2983-2984&amp;publication_year=2023&amp;author=Truhn%2CD&amp;author=Reis-Filho%2CJS&amp;author=Kather%2CJN\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR130\">Asgari, E. et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. NPJ Digit. Med. 8, 274 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41746-025-01670-7\" data-track-item_id=\"10.1038\/s41746-025-01670-7\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41746-025-01670-7\" aria-label=\"Article reference 130\" data-doi=\"10.1038\/s41746-025-01670-7\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40360677\" aria-label=\"PubMed reference 130\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12075489\" aria-label=\"PubMed Central reference 130\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 130\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=A%20framework%20to%20assess%20clinical%20safety%20and%20hallucination%20rates%20of%20LLMs%20for%20medical%20text%20summarisation&amp;journal=NPJ%20Digit.%20Med.&amp;doi=10.1038%2Fs41746-025-01670-7&amp;volume=8&amp;publication_year=2025&amp;author=Asgari%2CE\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR131\">Kim, Y. et al. Medical hallucinations in foundation models and their impact on healthcare. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2503.05777\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2503.05777\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2503.05777<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR132\">Bang, Y. et al. HalluLens: LLM hallucination benchmark. In Proc. 63rd Meeting of the Association for Computational Linguistics Vol. 1, 24128\u201324156 (ACL, 2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR133\">Soffer, S., Sorin, V., Nadkarni, G. N. &amp; Klang, E. Pitfalls of large language models in medical ethics reasoning. NPJ Digit. Med. 8, 461 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41746-025-01792-y\" data-track-item_id=\"10.1038\/s41746-025-01792-y\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41746-025-01792-y\" aria-label=\"Article reference 133\" data-doi=\"10.1038\/s41746-025-01792-y\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40696098\" aria-label=\"PubMed reference 133\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12284149\" aria-label=\"PubMed Central reference 133\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 133\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Pitfalls%20of%20large%20language%20models%20in%20medical%20ethics%20reasoning&amp;journal=NPJ%20Digit.%20Med.&amp;doi=10.1038%2Fs41746-025-01792-y&amp;volume=8&amp;publication_year=2025&amp;author=Soffer%2CS&amp;author=Sorin%2CV&amp;author=Nadkarni%2CGN&amp;author=Klang%2CE\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR134\">Xu, H. et al. Reducing tool hallucination via reliability alignment. In Proc. 42nd International Conference on Machine Learning 2799 (ACM, 2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR135\">Chung, P. et al. Verifying facts in patient care documents generated by large language models using electronic health records. NEJM AI 3, AIdbp2500418 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 135\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Verifying%20facts%20in%20patient%20care%20documents%20generated%20by%20large%20language%20models%20using%20electronic%20health%20records&amp;journal=NEJM%20AI&amp;volume=3&amp;publication_year=2025&amp;author=Chung%2CP\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR136\">World Medical Association. WMA Declaration of Helsinki\u2014Ethical Principles for Medical Research Involving Human Participants. WMA <a href=\"https:\/\/www.wma.net\/policies-post\/wma-declaration-of-helsinki\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/www.wma.net\/policies-post\/wma-declaration-of-helsinki\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/www.wma.net\/policies-post\/wma-declaration-of-helsinki<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR137\">Yu, K.-H., Healey, E., Leong, T.-Y., Kohane, I. S. &amp; Manrai, A. K. Medical artificial intelligence and human values. N. Engl. J. Med. 390, 1895\u20131904 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1056\/NEJMra2214183\" data-track-item_id=\"10.1056\/NEJMra2214183\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1056%2FNEJMra2214183\" aria-label=\"Article reference 137\" data-doi=\"10.1056\/NEJMra2214183\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=38810186\" aria-label=\"PubMed reference 137\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12425466\" aria-label=\"PubMed Central reference 137\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 137\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Medical%20artificial%20intelligence%20and%20human%20values&amp;journal=N.%20Engl.%20J.%20Med.&amp;doi=10.1056%2FNEJMra2214183&amp;volume=390&amp;pages=1895-1904&amp;publication_year=2024&amp;author=Yu%2CK-H&amp;author=Healey%2CE&amp;author=Leong%2CT-Y&amp;author=Kohane%2CIS&amp;author=Manrai%2CAK\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR138\">Mazeika, M. et al. Utility engineering: Analyzing and controlling emergent value systems in AIs. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2502.08640\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2502.08640\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2502.08640<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR139\">Greenblatt, R. et al. Alignment faking in large language models. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2412.14093\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2412.14093\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2412.14093<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR140\">Bedi, S. et al. Testing and evaluation of health care applications of large language models: A systematic review: A systematic review. JAMA 333, 319\u2013328 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1001\/jama.2024.21700\" data-track-item_id=\"10.1001\/jama.2024.21700\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1001%2Fjama.2024.21700\" aria-label=\"Article reference 140\" data-doi=\"10.1001\/jama.2024.21700\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 140\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Testing%20and%20evaluation%20of%20health%20care%20applications%20of%20large%20language%20models%3A%20A%20systematic%20review%3A%20A%20systematic%20review&amp;journal=JAMA&amp;doi=10.1001%2Fjama.2024.21700&amp;volume=333&amp;pages=319-328&amp;publication_year=2024&amp;author=Bedi%2CS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR141\">Zack, T. et al. Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care: a model evaluation study. Lancet Digit. Health 6, e12\u2013e22 (2024). This study highlights that LLMs can perpetuate racial and gender biases beyond evidence-based variation.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/S2589-7500(23)00225-X\" data-track-item_id=\"10.1016\/S2589-7500(23)00225-X\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2FS2589-7500%2823%2900225-X\" aria-label=\"Article reference 141\" data-doi=\"10.1016\/S2589-7500(23)00225-X\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB3sXis1Cnt73N\" aria-label=\"CAS reference 141\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=38123252\" aria-label=\"PubMed reference 141\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 141\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Assessing%20the%20potential%20of%20GPT-4%20to%20perpetuate%20racial%20and%20gender%20biases%20in%20health%20care%3A%20a%20model%20evaluation%20study&amp;journal=Lancet%20Digit.%20Health&amp;doi=10.1016%2FS2589-7500%2823%2900225-X&amp;volume=6&amp;pages=e12-e22&amp;publication_year=2024&amp;author=Zack%2CT\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR142\">Omar, M. et al. Sociodemographic biases in medical decision making by large language models. Nat. Med. 31, 1873\u20131881 (2025). This study assesses sociodemographic biases in medical LLMs, demonstrating differences in clinical decision-making that extend beyond evidence-based variation.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41591-025-03626-6\" data-track-item_id=\"10.1038\/s41591-025-03626-6\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41591-025-03626-6\" aria-label=\"Article reference 142\" data-doi=\"10.1038\/s41591-025-03626-6\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXhsVSqsrbO\" aria-label=\"CAS reference 142\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40195448\" aria-label=\"PubMed reference 142\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 142\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Sociodemographic%20biases%20in%20medical%20decision%20making%20by%20large%20language%20models&amp;journal=Nat.%20Med.&amp;doi=10.1038%2Fs41591-025-03626-6&amp;volume=31&amp;pages=1873-1881&amp;publication_year=2025&amp;author=Omar%2CM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR143\">Gruber, V.-E. et al. A women\u2019s health benchmark for large language models. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2512.17028\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2512.17028\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2512.17028<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR144\">Yang, J., Soltan, A. A. S., Eyre, D. W., Yang, Y. &amp; Clifton, D. A. An adversarial training framework for mitigating algorithmic biases in clinical machine learning. NPJ Digit. Med. 6, 55 (2023).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41746-023-00805-y\" data-track-item_id=\"10.1038\/s41746-023-00805-y\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41746-023-00805-y\" aria-label=\"Article reference 144\" data-doi=\"10.1038\/s41746-023-00805-y\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=36991077\" aria-label=\"PubMed reference 144\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC10050816\" aria-label=\"PubMed Central reference 144\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 144\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=An%20adversarial%20training%20framework%20for%20mitigating%20algorithmic%20biases%20in%20clinical%20machine%20learning&amp;journal=NPJ%20Digit.%20Med.&amp;doi=10.1038%2Fs41746-023-00805-y&amp;volume=6&amp;publication_year=2023&amp;author=Yang%2CJ&amp;author=Soltan%2CAAS&amp;author=Eyre%2CDW&amp;author=Yang%2CY&amp;author=Clifton%2CDA\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR145\">Gichoya, J. W. et al. AI pitfalls and what not to do: mitigating bias in AI. Br. J. Radiol. 96, 20230023 (2023).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1259\/bjr.20230023\" data-track-item_id=\"10.1259\/bjr.20230023\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1259%2Fbjr.20230023\" aria-label=\"Article reference 145\" data-doi=\"10.1259\/bjr.20230023\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=37698583\" aria-label=\"PubMed reference 145\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC10546443\" aria-label=\"PubMed Central reference 145\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 145\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=AI%20pitfalls%20and%20what%20not%20to%20do%3A%20mitigating%20bias%20in%20AI&amp;journal=Br.%20J.%20Radiol.&amp;doi=10.1259%2Fbjr.20230023&amp;volume=96&amp;publication_year=2023&amp;author=Gichoya%2CJW\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR146\">Ktena, I. et al. Generative models improve fairness of medical classifiers under distribution shifts. Nat. Med. 30, 1166\u20131173 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41591-024-02838-6\" data-track-item_id=\"10.1038\/s41591-024-02838-6\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41591-024-02838-6\" aria-label=\"Article reference 146\" data-doi=\"10.1038\/s41591-024-02838-6\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2cXnvF2ltb4%3D\" aria-label=\"CAS reference 146\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=38600282\" aria-label=\"PubMed reference 146\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11031395\" aria-label=\"PubMed Central reference 146\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 146\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Generative%20models%20improve%20fairness%20of%20medical%20classifiers%20under%20distribution%20shifts&amp;journal=Nat.%20Med.&amp;doi=10.1038%2Fs41591-024-02838-6&amp;volume=30&amp;pages=1166-1173&amp;publication_year=2024&amp;author=Ktena%2CI\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR147\">Xu, Z. Mitigating social bias in large language models: a multi-objective approach within a multi-agent framework. In Proc. 39th Conf. AAAI Artificial Intelligence 25587\u201325579 (AAAI, 2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR148\">Binz, M. et al. A foundation model to predict and capture human cognition. Nature 644, 1002\u20131009 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41586-025-09215-4\" data-track-item_id=\"10.1038\/s41586-025-09215-4\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41586-025-09215-4\" aria-label=\"Article reference 148\" data-doi=\"10.1038\/s41586-025-09215-4\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2025Natur.644.1002B\" aria-label=\"ADS reference 148\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXhvV2nsr7P\" aria-label=\"CAS reference 148\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40604288\" aria-label=\"PubMed reference 148\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12390832\" aria-label=\"PubMed Central reference 148\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 148\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=A%20foundation%20model%20to%20predict%20and%20capture%20human%20cognition&amp;journal=Nature&amp;doi=10.1038%2Fs41586-025-09215-4&amp;volume=644&amp;pages=1002-1009&amp;publication_year=2025&amp;author=Binz%2CM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR149\">Mendu, S. K., Yenala, H., Gulati, A., Kumar, S. &amp; Agrawal, P. Towards safer pretraining: analyzing and filtering harmful content in webscale datasets for responsible LLMs. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2505.02009\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2505.02009\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2505.02009<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR150\">Griot, M., Hemptinne, C., Vanderdonckt, J. &amp; Yuksel, D. Large language models lack essential metacognition for reliable medical reasoning. Nat. Commun. 16, 642 (2025). This study reveals a gap between benchmark performance and metacognitive awareness in LLMs.<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41467-024-55628-6\" data-track-item_id=\"10.1038\/s41467-024-55628-6\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41467-024-55628-6\" aria-label=\"Article reference 150\" data-doi=\"10.1038\/s41467-024-55628-6\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"ads reference\" data-track-action=\"ads reference\" href=\"http:\/\/adsabs.harvard.edu\/cgi-bin\/nph-data_query?link_type=ABSTRACT&amp;bibcode=2025NatCo..16..642G\" aria-label=\"ADS reference 150\" target=\"_blank\">ADS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXhsFGgur8%3D\" aria-label=\"CAS reference 150\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39809759\" aria-label=\"PubMed reference 150\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11733150\" aria-label=\"PubMed Central reference 150\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 150\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Large%20language%20models%20lack%20essential%20metacognition%20for%20reliable%20medical%20reasoning&amp;journal=Nat.%20Commun.&amp;doi=10.1038%2Fs41467-024-55628-6&amp;volume=16&amp;publication_year=2025&amp;author=Griot%2CM&amp;author=Hemptinne%2CC&amp;author=Vanderdonckt%2CJ&amp;author=Yuksel%2CD\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR151\">van Buchem, M. M. et al. Impact of a digital scribe system on clinical documentation time and quality: Usability study. JMIR AI 3, e60020 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.2196\/60020\" data-track-item_id=\"10.2196\/60020\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.2196%2F60020\" aria-label=\"Article reference 151\" data-doi=\"10.2196\/60020\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39312397\" aria-label=\"PubMed reference 151\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11459111\" aria-label=\"PubMed Central reference 151\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 151\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Impact%20of%20a%20digital%20scribe%20system%20on%20clinical%20documentation%20time%20and%20quality%3A%20Usability%20study&amp;journal=JMIR%20AI&amp;doi=10.2196%2F60020&amp;volume=3&amp;publication_year=2024&amp;author=Buchem%2CMM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR152\">Ma, S. P. et al. Ambient artificial intelligence scribes: utilization and impact on documentation time. J. Am. Med. Inform. Assoc. 32, 381\u2013385 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1093\/jamia\/ocae304\" data-track-item_id=\"10.1093\/jamia\/ocae304\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1093%2Fjamia%2Focae304\" aria-label=\"Article reference 152\" data-doi=\"10.1093\/jamia\/ocae304\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39688515\" aria-label=\"PubMed reference 152\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11756633\" aria-label=\"PubMed Central reference 152\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 152\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Ambient%20artificial%20intelligence%20scribes%3A%20utilization%20and%20impact%20on%20documentation%20time&amp;journal=J.%20Am.%20Med.%20Inform.%20Assoc.&amp;doi=10.1093%2Fjamia%2Focae304&amp;volume=32&amp;pages=381-385&amp;publication_year=2025&amp;author=Ma%2CSP\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR153\">You, J. G. et al. Ambient documentation technology in clinician experience of documentation burden and burnout. JAMA Netw. Open 8, e2528056 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1001\/jamanetworkopen.2025.28056\" data-track-item_id=\"10.1001\/jamanetworkopen.2025.28056\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1001%2Fjamanetworkopen.2025.28056\" aria-label=\"Article reference 153\" data-doi=\"10.1001\/jamanetworkopen.2025.28056\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40839265\" aria-label=\"PubMed reference 153\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12371510\" aria-label=\"PubMed Central reference 153\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 153\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Ambient%20documentation%20technology%20in%20clinician%20experience%20of%20documentation%20burden%20and%20burnout&amp;journal=JAMA%20Netw.%20Open&amp;doi=10.1001%2Fjamanetworkopen.2025.28056&amp;volume=8&amp;publication_year=2025&amp;author=You%2CJG\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR154\">Blease, C. R., Locher, C., Gaab, J., H\u00e4gglund, M. &amp; Mandl, K. D. Generative artificial intelligence in primary care: an online survey of UK general practitioners. BMJ Health Care Inform. 31, e101102 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1136\/bmjhci-2024-101102\" data-track-item_id=\"10.1136\/bmjhci-2024-101102\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1136%2Fbmjhci-2024-101102\" aria-label=\"Article reference 154\" data-doi=\"10.1136\/bmjhci-2024-101102\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39288998\" aria-label=\"PubMed reference 154\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11429366\" aria-label=\"PubMed Central reference 154\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 154\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Generative%20artificial%20intelligence%20in%20primary%20care%3A%20an%20online%20survey%20of%20UK%20general%20practitioners&amp;journal=BMJ%20Health%20Care%20Inform.&amp;doi=10.1136%2Fbmjhci-2024-101102&amp;volume=31&amp;publication_year=2024&amp;author=Blease%2CCR&amp;author=Locher%2CC&amp;author=Gaab%2CJ&amp;author=H%C3%A4gglund%2CM&amp;author=Mandl%2CKD\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR155\">Eppler, M. et al. Awareness and use of ChatGPT and large language models: a prospective cross-sectional global survey in urology. Eur. Urol. 85, 146\u2013153 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/j.eururo.2023.10.014\" data-track-item_id=\"10.1016\/j.eururo.2023.10.014\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2Fj.eururo.2023.10.014\" aria-label=\"Article reference 155\" data-doi=\"10.1016\/j.eururo.2023.10.014\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=37926642\" aria-label=\"PubMed reference 155\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 155\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Awareness%20and%20use%20of%20ChatGPT%20and%20large%20language%20models%3A%20a%20prospective%20cross-sectional%20global%20survey%20in%20urology&amp;journal=Eur.%20Urol.&amp;doi=10.1016%2Fj.eururo.2023.10.014&amp;volume=85&amp;pages=146-153&amp;publication_year=2024&amp;author=Eppler%2CM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR156\">Poon, E. G., Lemak, C. H., Rojas, J. C., Guptill, J. &amp; Classen, D. Adoption of artificial intelligence in healthcare: survey of health system priorities, successes, and challenges. J. Am. Med. Inform. Assoc. 32, 1093\u20131100 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1093\/jamia\/ocaf065\" data-track-item_id=\"10.1093\/jamia\/ocaf065\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1093%2Fjamia%2Focaf065\" aria-label=\"Article reference 156\" data-doi=\"10.1093\/jamia\/ocaf065\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40323320\" aria-label=\"PubMed reference 156\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12202002\" aria-label=\"PubMed Central reference 156\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 156\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Adoption%20of%20artificial%20intelligence%20in%20healthcare%3A%20survey%20of%20health%20system%20priorities%2C%20successes%2C%20and%20challenges&amp;journal=J.%20Am.%20Med.%20Inform.%20Assoc.&amp;doi=10.1093%2Fjamia%2Focaf065&amp;volume=32&amp;pages=1093-1100&amp;publication_year=2025&amp;author=Poon%2CEG&amp;author=Lemak%2CCH&amp;author=Rojas%2CJC&amp;author=Guptill%2CJ&amp;author=Classen%2CD\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR157\">Egli, S. B., Arpagaus, A., Amacher, S. A., Hunziker, S. &amp; Bassetti, S. Use, knowledge and perception of large language models in clinical practice: a cross-sectional mixed-methods survey among clinicians in Switzerland. BMJ Health Care Inform. 32, e101470 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1136\/bmjhci-2025-101470\" data-track-item_id=\"10.1136\/bmjhci-2025-101470\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1136%2Fbmjhci-2025-101470\" aria-label=\"Article reference 157\" data-doi=\"10.1136\/bmjhci-2025-101470\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40983363\" aria-label=\"PubMed reference 157\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12458741\" aria-label=\"PubMed Central reference 157\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 157\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Use%2C%20knowledge%20and%20perception%20of%20large%20language%20models%20in%20clinical%20practice%3A%20a%20cross-sectional%20mixed-methods%20survey%20among%20clinicians%20in%20Switzerland&amp;journal=BMJ%20Health%20Care%20Inform.&amp;doi=10.1136%2Fbmjhci-2025-101470&amp;volume=32&amp;publication_year=2025&amp;author=Egli%2CSB&amp;author=Arpagaus%2CA&amp;author=Amacher%2CSA&amp;author=Hunziker%2CS&amp;author=Bassetti%2CS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR158\">Castiblanco Jimenez, I. A., Gomez Acevedo, J. S., Marcolin, F., Vezzetti, E. &amp; Moos, S. Towards an integrated framework to measure user engagement with interactive or physical products. Int. J. Interact. Des. Manuf. 17, 45\u201367 (2023).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"noopener nofollow\" data-track-label=\"10.1007\/s12008-022-01087-6\" data-track-item_id=\"10.1007\/s12008-022-01087-6\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/link.springer.com\/doi\/10.1007\/s12008-022-01087-6\" aria-label=\"Article reference 158\" data-doi=\"10.1007\/s12008-022-01087-6\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 158\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Towards%20an%20integrated%20framework%20to%20measure%20user%20engagement%20with%20interactive%20or%20physical%20products&amp;journal=Int.%20J.%20Interact.%20Des.%20Manuf.&amp;doi=10.1007%2Fs12008-022-01087-6&amp;volume=17&amp;pages=45-67&amp;publication_year=2023&amp;author=Castiblanco%20Jimenez%2CIA&amp;author=Gomez%20Acevedo%2CJS&amp;author=Marcolin%2CF&amp;author=Vezzetti%2CE&amp;author=Moos%2CS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR159\">Chang, C. T. et al. Red teaming ChatGPT in medicine to yield real-world insights on model behavior. NPJ Digit. Med. 8, 149 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41746-025-01542-0\" data-track-item_id=\"10.1038\/s41746-025-01542-0\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41746-025-01542-0\" aria-label=\"Article reference 159\" data-doi=\"10.1038\/s41746-025-01542-0\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40055532\" aria-label=\"PubMed reference 159\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11889229\" aria-label=\"PubMed Central reference 159\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 159\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Red%20teaming%20ChatGPT%20in%20medicine%20to%20yield%20real-world%20insights%20on%20model%20behavior&amp;journal=NPJ%20Digit.%20Med.&amp;doi=10.1038%2Fs41746-025-01542-0&amp;volume=8&amp;publication_year=2025&amp;author=Chang%2CCT\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR160\">Artsi, Y. et al. Large language models in real-world clinical workflows: a systematic review of applications and implementation. Front. Digit. Health 7, 1659134 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.3389\/fdgth.2025.1659134\" data-track-item_id=\"10.3389\/fdgth.2025.1659134\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.3389%2Ffdgth.2025.1659134\" aria-label=\"Article reference 160\" data-doi=\"10.3389\/fdgth.2025.1659134\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=41098649\" aria-label=\"PubMed reference 160\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12519456\" aria-label=\"PubMed Central reference 160\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 160\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Large%20language%20models%20in%20real-world%20clinical%20workflows%3A%20a%20systematic%20review%20of%20applications%20and%20implementation&amp;journal=Front.%20Digit.%20Health&amp;doi=10.3389%2Ffdgth.2025.1659134&amp;volume=7&amp;publication_year=2025&amp;author=Artsi%2CY\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR161\">Bondi-Kelly, E. et al. Taking off with AI: lessons from aviation for healthcare. In Proc. 3rd ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization <a href=\"https:\/\/doi.org\/10.1145\/3617694.3623224\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.1145\/3617694.3623224\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.1145\/3617694.3623224<\/a> (ACM, 2023).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR162\">Kolbinger, F. R. &amp; Kather, J. N. Adaptive validation strategies for real-world clinical artificial intelligence. Nat. Comput. Sci. 5, 980\u2013986 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s43588-025-00901-x\" data-track-item_id=\"10.1038\/s43588-025-00901-x\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs43588-025-00901-x\" aria-label=\"Article reference 162\" data-doi=\"10.1038\/s43588-025-00901-x\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=41249673\" aria-label=\"PubMed reference 162\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 162\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Adaptive%20validation%20strategies%20for%20real-world%20clinical%20artificial%20intelligence&amp;journal=Nat.%20Comput.%20Sci.&amp;doi=10.1038%2Fs43588-025-00901-x&amp;volume=5&amp;pages=980-986&amp;publication_year=2025&amp;author=Kolbinger%2CFR&amp;author=Kather%2CJN\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR163\">Crowe, B. et al. Recommendations for clinicians, technologists, and healthcare organizations on the use of generative artificial intelligence in medicine: a position statement from the Society of General Internal Medicine. J. Gen. Intern. Med. 40, 694\u2013702 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"noopener nofollow\" data-track-label=\"10.1007\/s11606-024-09102-0\" data-track-item_id=\"10.1007\/s11606-024-09102-0\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/link.springer.com\/doi\/10.1007\/s11606-024-09102-0\" aria-label=\"Article reference 163\" data-doi=\"10.1007\/s11606-024-09102-0\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39531100\" aria-label=\"PubMed reference 163\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 163\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Recommendations%20for%20clinicians%2C%20technologists%2C%20and%20healthcare%20organizations%20on%20the%20use%20of%20generative%20artificial%20intelligence%20in%20medicine%3A%20a%20position%20statement%20from%20the%20Society%20of%20General%20Internal%20Medicine&amp;journal=J.%20Gen.%20Intern.%20Med.&amp;doi=10.1007%2Fs11606-024-09102-0&amp;volume=40&amp;pages=694-702&amp;publication_year=2025&amp;author=Crowe%2CB\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR164\">Zhou, A. et al. AutoRedTeamer: autonomous red teaming with lifelong attack integration. NeurIPS 2025 Conference <a href=\"https:\/\/openreview.net\/forum?id=xQH4lDLIC0\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/openreview.net\/forum?id=xQH4lDLIC0\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/openreview.net\/forum?id=xQH4lDLIC0<\/a> (NeurIPS, 2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR165\">Bastani, H. et al. Generative AI without guardrails can harm learning: evidence from high school mathematics. Proc. Natl Acad. Sci. USA 122, e2422633122 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1073\/pnas.2422633122\" data-track-item_id=\"10.1073\/pnas.2422633122\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1073%2Fpnas.2422633122\" aria-label=\"Article reference 165\" data-doi=\"10.1073\/pnas.2422633122\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXitFSmtLvN\" aria-label=\"CAS reference 165\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40560616\" aria-label=\"PubMed reference 165\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12232635\" aria-label=\"PubMed Central reference 165\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 165\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Generative%20AI%20without%20guardrails%20can%20harm%20learning%3A%20evidence%20from%20high%20school%20mathematics&amp;journal=Proc.%20Natl%20Acad.%20Sci.%20USA&amp;doi=10.1073%2Fpnas.2422633122&amp;volume=122&amp;publication_year=2025&amp;author=Bastani%2CH\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR166\">Mollick, E. R. et al. AI Agents and Education: Simulated Practice at Scale. The Wharton School Research Paper <a href=\"https:\/\/doi.org\/10.2139\/ssrn.4871171\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.2139\/ssrn.4871171\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.2139\/ssrn.4871171<\/a> (SSRN, 2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR167\">Kosmyna, N. et al. Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2506.08872\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2506.08872\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2506.08872<\/a> (2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR168\">Hoffmann, M., Boysel, S., Nagle, F., Peng, S. &amp; Xu, K. Generative AI and the nature of work. Harvard Business School Working Paper 25-021 <a href=\"https:\/\/doi.org\/10.2139\/ssrn.5007084\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.2139\/ssrn.5007084\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.2139\/ssrn.5007084<\/a> (SSRN, 2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR169\">Wekenborg, M. K., Gilbert, S. &amp; Kather, J. N. Examining human\u2013AI interaction in real-world healthcare beyond the laboratory. NPJ Digit. Med. 8, 169 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41746-025-01559-5\" data-track-item_id=\"10.1038\/s41746-025-01559-5\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41746-025-01559-5\" aria-label=\"Article reference 169\" data-doi=\"10.1038\/s41746-025-01559-5\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40108434\" aria-label=\"PubMed reference 169\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11923224\" aria-label=\"PubMed Central reference 169\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 169\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Examining%20human%E2%80%93AI%20interaction%20in%20real-world%20healthcare%20beyond%20the%20laboratory&amp;journal=NPJ%20Digit.%20Med.&amp;doi=10.1038%2Fs41746-025-01559-5&amp;volume=8&amp;publication_year=2025&amp;author=Wekenborg%2CMK&amp;author=Gilbert%2CS&amp;author=Kather%2CJN\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR170\">Berzin, T. M. &amp; Topol, E. J. Preserving clinical skills in the age of AI assistance. Lancet 406, 1719 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/S0140-6736(25)02075-6\" data-track-item_id=\"10.1016\/S0140-6736(25)02075-6\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2FS0140-6736%2825%2902075-6\" aria-label=\"Article reference 170\" data-doi=\"10.1016\/S0140-6736(25)02075-6\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=41109709\" aria-label=\"PubMed reference 170\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 170\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Preserving%20clinical%20skills%20in%20the%20age%20of%20AI%20assistance&amp;journal=Lancet&amp;doi=10.1016%2FS0140-6736%2825%2902075-6&amp;volume=406&amp;publication_year=2025&amp;author=Berzin%2CTM&amp;author=Topol%2CEJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR171\">Budzy\u0144, K. et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterol. Hepatol. 10, 896\u2013903 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/S2468-1253(25)00133-5\" data-track-item_id=\"10.1016\/S2468-1253(25)00133-5\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2FS2468-1253%2825%2900133-5\" aria-label=\"Article reference 171\" data-doi=\"10.1016\/S2468-1253(25)00133-5\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40816301\" aria-label=\"PubMed reference 171\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 171\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Endoscopist%20deskilling%20risk%20after%20exposure%20to%20artificial%20intelligence%20in%20colonoscopy%3A%20a%20multicentre%2C%20observational%20study&amp;journal=Lancet%20Gastroenterol.%20Hepatol.&amp;doi=10.1016%2FS2468-1253%2825%2900133-5&amp;volume=10&amp;pages=896-903&amp;publication_year=2025&amp;author=Budzy%C5%84%2CK\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR172\">Huo, W., Li, Q., Liang, B., Wang, Y. &amp; Li, X. When healthcare professionals use AI: exploring work well-being through psychological needs satisfaction and job complexity. Behav. Sci. 15, 88 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.3390\/bs15010088\" data-track-item_id=\"10.3390\/bs15010088\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.3390%2Fbs15010088\" aria-label=\"Article reference 172\" data-doi=\"10.3390\/bs15010088\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39851892\" aria-label=\"PubMed reference 172\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11761562\" aria-label=\"PubMed Central reference 172\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 172\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=When%20healthcare%20professionals%20use%20AI%3A%20exploring%20work%20well-being%20through%20psychological%20needs%20satisfaction%20and%20job%20complexity&amp;journal=Behav.%20Sci.&amp;doi=10.3390%2Fbs15010088&amp;volume=15&amp;publication_year=2025&amp;author=Huo%2CW&amp;author=Li%2CQ&amp;author=Liang%2CB&amp;author=Wang%2CY&amp;author=Li%2CX\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR173\">Rafailov, R. et al. Direct preference optimization: your language model is secretly a reward model. NeurIPS 2023 Conference <a href=\"https:\/\/neurips.cc\/virtual\/2023\/poster\/72164\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/neurips.cc\/virtual\/2023\/poster\/72164\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/neurips.cc\/virtual\/2023\/poster\/72164<\/a> (NeurIPS, 2023).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR174\">Zou, A. et al. Representation engineering: a top-down approach to AI transparency <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2310.01405\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2310.01405\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2310.01405<\/a> (2023).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR175\">Zou, A. et al. Improving alignment and robustness with circuit breakers. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2406.04313\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2406.04313\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2406.04313<\/a> (2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR176\">Inan, H. et al. Llama Guard: LLM-based input\u2013output safeguard for Human\u2013AI conversations. Preprint at <a href=\"https:\/\/doi.org\/10.48550\/arXiv.2312.06674\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"10.48550\/arXiv.2312.06674\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/doi.org\/10.48550\/arXiv.2312.06674<\/a> (2023).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR177\">Rebedea, T., Dinu, R., Sreedhar, M. N., Parisien, C. &amp; Cohen, J. NeMo guardrails: a toolkit for controllable and safe LLM applications with programmable rails. In Proc. 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations 431\u2013445 (Association for Computational Linguistics, 2023).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR178\">Alaa, A. et al. Position: medical large language model benchmarks should prioritize construct validity Oral. In Proc. International Conference on Machine Learning 2025 <a href=\"https:\/\/icml.cc\/virtual\/2025\/oral\/40130\" data-track=\"click_references\" data-track-action=\"external reference\" data-track-value=\"external reference\" data-track-label=\"https:\/\/icml.cc\/virtual\/2025\/oral\/40130\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/icml.cc\/virtual\/2025\/oral\/40130<\/a> (ICML, 2025).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR179\">Gallifant, J. &amp; Bitterman, D. S. Humanity\u2019s next medical exam: preparing to evaluate superhuman systems. NEJM AI 2, AIe2501008 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1056\/AIe2501008\" data-track-item_id=\"10.1056\/AIe2501008\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1056%2FAIe2501008\" aria-label=\"Article reference 179\" data-doi=\"10.1056\/AIe2501008\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 179\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Humanity%E2%80%99s%20next%20medical%20exam%3A%20preparing%20to%20evaluate%20superhuman%20systems&amp;journal=NEJM%20AI&amp;doi=10.1056%2FAIe2501008&amp;volume=2&amp;publication_year=2025&amp;author=Gallifant%2CJ&amp;author=Bitterman%2CDS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR180\">Freyer, O. et al. Consideration of cybersecurity risks in the benefit-risk analysis of medical devices: scoping review. J. Med. Internet Res. 26, e65528 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.2196\/65528\" data-track-item_id=\"10.2196\/65528\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.2196%2F65528\" aria-label=\"Article reference 180\" data-doi=\"10.2196\/65528\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39718821\" aria-label=\"PubMed reference 180\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC11707448\" aria-label=\"PubMed Central reference 180\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 180\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Consideration%20of%20cybersecurity%20risks%20in%20the%20benefit-risk%20analysis%20of%20medical%20devices%3A%20scoping%20review&amp;journal=J.%20Med.%20Internet%20Res.&amp;doi=10.2196%2F65528&amp;volume=26&amp;publication_year=2024&amp;author=Freyer%2CO\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR181\">Ostermann, M. et al. Cybersecurity requirements for medical devices in the EU and US\u2014a comparison and gap analysis of the MDCG 2019-16 and FDA premarket cybersecurity guidance. Comput. Struct. Biotechnol. J. 28, 259\u2013266 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/j.csbj.2025.07.024\" data-track-item_id=\"10.1016\/j.csbj.2025.07.024\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2Fj.csbj.2025.07.024\" aria-label=\"Article reference 181\" data-doi=\"10.1016\/j.csbj.2025.07.024\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40727673\" aria-label=\"PubMed reference 181\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12301760\" aria-label=\"PubMed Central reference 181\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 181\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Cybersecurity%20requirements%20for%20medical%20devices%20in%20the%20EU%20and%20US%E2%80%94a%20comparison%20and%20gap%20analysis%20of%20the%20MDCG%202019-16%20and%20FDA%20premarket%20cybersecurity%20guidance&amp;journal=Comput.%20Struct.%20Biotechnol.%20J.&amp;doi=10.1016%2Fj.csbj.2025.07.024&amp;volume=28&amp;pages=259-266&amp;publication_year=2025&amp;author=Ostermann%2CM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR182\">Moberly, T. Doctors must stop using unregistered AI scribe tools, says NHS England. Brit. Med. J. 389, r1302 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1136\/bmj.r1302\" data-track-item_id=\"10.1136\/bmj.r1302\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1136%2Fbmj.r1302\" aria-label=\"Article reference 182\" data-doi=\"10.1136\/bmj.r1302\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40550588\" aria-label=\"PubMed reference 182\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 182\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Doctors%20must%20stop%20using%20unregistered%20AI%20scribe%20tools%2C%20says%20NHS%20England&amp;journal=Brit.%20Med.%20J.&amp;doi=10.1136%2Fbmj.r1302&amp;volume=389&amp;publication_year=2025&amp;author=Moberly%2CT\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR183\">Freyer, O., Wiest, I. C., Kather, J. N. &amp; Gilbert, S. A future role for health applications of large language models depends on regulators enforcing safety standards. Lancet Digit. Health 6, e662\u2013e672 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1016\/S2589-7500(24)00124-9\" data-track-item_id=\"10.1016\/S2589-7500(24)00124-9\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1016%2FS2589-7500%2824%2900124-9\" aria-label=\"Article reference 183\" data-doi=\"10.1016\/S2589-7500(24)00124-9\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2cXhvVWhtLjM\" aria-label=\"CAS reference 183\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=39179311\" aria-label=\"PubMed reference 183\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 183\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=A%20future%20role%20for%20health%20applications%20of%20large%20language%20models%20depends%20on%20regulators%20enforcing%20safety%20standards&amp;journal=Lancet%20Digit.%20Health&amp;doi=10.1016%2FS2589-7500%2824%2900124-9&amp;volume=6&amp;pages=e662-e672&amp;publication_year=2024&amp;author=Freyer%2CO&amp;author=Wiest%2CIC&amp;author=Kather%2CJN&amp;author=Gilbert%2CS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR184\">Biasin, E., Kamenja\u0161evi\u0107, E. &amp; Ludvigsen, K. R. in Research Handbook on Health, AI and the Law (eds Solaiman, B. &amp; Cohen, I. G.) 57\u201374 (Edward Elgar, 2024).<\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR185\">Freyer, O., Jayabalan, S., Kather, J. N. &amp; Gilbert, S. Overcoming regulatory barriers to the implementation of AI agents in healthcare. Nat. Med. 31, 3239\u20133243 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s41591-025-03841-1\" data-track-item_id=\"10.1038\/s41591-025-03841-1\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs41591-025-03841-1\" aria-label=\"Article reference 185\" data-doi=\"10.1038\/s41591-025-03841-1\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"cas reference\" data-track-action=\"cas reference\" href=\"https:\/\/www.nature.com\/articles\/cas-redirect\/1:CAS:528:DC%2BB2MXhvFCis7%2FJ\" aria-label=\"CAS reference 185\" target=\"_blank\">CAS<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40681675\" aria-label=\"PubMed reference 185\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 185\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Overcoming%20regulatory%20barriers%20to%20the%20implementation%20of%20AI%20agents%20in%20healthcare&amp;journal=Nat.%20Med.&amp;doi=10.1038%2Fs41591-025-03841-1&amp;volume=31&amp;pages=3239-3243&amp;publication_year=2025&amp;author=Freyer%2CO&amp;author=Jayabalan%2CS&amp;author=Kather%2CJN&amp;author=Gilbert%2CS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR186\">Mathias, R., Schonfelder, A., Welzel, C. &amp; Gilbert, S. Letter to the editor on \u2018From concept to clinic: living labs and regulatory sandboxes for health system digitalization and the integration of innovative devices into clinical workflows\u2019. IEEE J. Transl. Eng. Health Med. 13, 214\u2013215 (2025).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1109\/JTEHM.2025.3557508\" data-track-item_id=\"10.1109\/JTEHM.2025.3557508\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1109%2FJTEHM.2025.3557508\" aria-label=\"Article reference 186\" data-doi=\"10.1109\/JTEHM.2025.3557508\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=40657527\" aria-label=\"PubMed reference 186\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC12251058\" aria-label=\"PubMed Central reference 186\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 186\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Letter%20to%20the%20editor%20on%20%E2%80%98From%20concept%20to%20clinic%3A%20living%20labs%20and%20regulatory%20sandboxes%20for%20health%20system%20digitalization%20and%20the%20integration%20of%20innovative%20devices%20into%20clinical%20workflows%E2%80%99&amp;journal=IEEE%20J.%20Transl.%20Eng.%20Health%20Med.&amp;doi=10.1109%2FJTEHM.2025.3557508&amp;volume=13&amp;pages=214-215&amp;publication_year=2025&amp;author=Mathias%2CR&amp;author=Schonfelder%2CA&amp;author=Welzel%2CC&amp;author=Gilbert%2CS\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR187\">Adler-Milstein, J. et al. Electronic health record adoption in US hospitals: the emergence of a digital \u2018advanced use\u2019 divide. J. Am. Med. Inform. Assoc. 24, 1142\u20131148 (2017).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1093\/jamia\/ocx080\" data-track-item_id=\"10.1093\/jamia\/ocx080\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1093%2Fjamia%2Focx080\" aria-label=\"Article reference 187\" data-doi=\"10.1093\/jamia\/ocx080\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed reference\" data-track-action=\"pubmed reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/entrez\/query.fcgi?cmd=Retrieve&amp;db=PubMed&amp;dopt=Abstract&amp;list_uids=29016973\" aria-label=\"PubMed reference 187\" target=\"_blank\">PubMed<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"link\" data-track-item_id=\"link\" data-track-value=\"pubmed central reference\" data-track-action=\"pubmed central reference\" href=\"http:\/\/www.ncbi.nlm.nih.gov\/pmc\/articles\/PMC7651985\" aria-label=\"PubMed Central reference 187\" target=\"_blank\">PubMed Central<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 187\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Electronic%20health%20record%20adoption%20in%20US%20hospitals%3A%20the%20emergence%20of%20a%20digital%20%E2%80%98advanced%20use%E2%80%99%20divide&amp;journal=J.%20Am.%20Med.%20Inform.%20Assoc.&amp;doi=10.1093%2Fjamia%2Focx080&amp;volume=24&amp;pages=1142-1148&amp;publication_year=2017&amp;author=Adler-Milstein%2CJ\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR188\">Hwang, Y.-M., Ng, M. Y., Pillai, M., Sahai, M. P. &amp; Hernandez-Boussard, T. The landscape of AI implementation in US hospitals. Nat. Health 1, 99\u2013112 (2026).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1038\/s44360-025-00016-7\" data-track-item_id=\"10.1038\/s44360-025-00016-7\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1038%2Fs44360-025-00016-7\" aria-label=\"Article reference 188\" data-doi=\"10.1038\/s44360-025-00016-7\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 188\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=The%20landscape%20of%20AI%20implementation%20in%20US%20hospitals&amp;journal=Nat.%20Health&amp;doi=10.1038%2Fs44360-025-00016-7&amp;volume=1&amp;pages=99-112&amp;publication_year=2026&amp;author=Hwang%2CY-M&amp;author=Ng%2CMY&amp;author=Pillai%2CM&amp;author=Sahai%2CMP&amp;author=Hernandez-Boussard%2CT\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR189\">Shanahan, M. Talking about large language models. Commun. ACM 67, 68\u201379 (2024).<\/p>\n<p class=\"c-article-references__links u-hide-print\"><a data-track=\"click_references\" rel=\"nofollow noopener\" data-track-label=\"10.1145\/3624724\" data-track-item_id=\"10.1145\/3624724\" data-track-value=\"article reference\" data-track-action=\"article reference\" href=\"https:\/\/doi.org\/10.1145%2F3624724\" aria-label=\"Article reference 189\" data-doi=\"10.1145\/3624724\" target=\"_blank\">Article<\/a>\u00a0<br \/>\n    <a data-track=\"click_references\" data-track-action=\"google scholar reference\" data-track-value=\"google scholar reference\" data-track-label=\"link\" data-track-item_id=\"link\" rel=\"nofollow noopener\" aria-label=\"Google Scholar reference 189\" href=\"http:\/\/scholar.google.com\/scholar_lookup?&amp;title=Talking%20about%20large%20language%20models&amp;journal=Commun.%20ACM&amp;doi=10.1145%2F3624724&amp;volume=67&amp;pages=68-79&amp;publication_year=2024&amp;author=Shanahan%2CM\" target=\"_blank\"><br \/>\n                    Google Scholar<\/a>\u00a0\n                <\/p>\n<p class=\"c-article-references__text\" id=\"ref-CR190\">Ferber, D. et al. Towards autonomous medical artificial intelligence agents. Nature 655, 1282\u20131291 (2026).<\/p>\n","protected":false},"excerpt":{"rendered":"Gommers, J. et al. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without&hellip;\n","protected":false},"author":2,"featured_media":819474,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[59],"tags":[6126,19944,10213,97,252,253,1159,17623,1160,79,19241],"class_list":["post-819473","post","type-post","status-publish","format-standard","has-post-thumbnail","category-health-care","tag-computational-science","tag-computer-science","tag-diagnosis","tag-health","tag-health-care","tag-healthcare","tag-humanities-and-social-sciences","tag-medical-ethics","tag-multidisciplinary","tag-science","tag-translational-research"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/819473","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/comments?post=819473"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/posts\/819473\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media\/819474"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/media?parent=819473"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/categories?post=819473"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/us\/wp-json\/wp\/v2\/tags?post=819473"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}