Jaech, A. et al. OpenAI o1 system card. Preprint at https://arxiv.org/abs/2412.16720 (2024).
Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 633–638 (2025).
Trinh, T. H., Wu, Y., Le, Q. V., He, H. & Luong, T. Solving olympiad geometry without human demonstrations. Nature 625, 476–482 (2024).
Rajpurkar, P., Chen, E., Banerjee, O. & Topol, E. J. AI in health and medicine. Nat. Med. 28, 31–38 (2022).
Mamede, S. et al. Immunising’physicians against availability bias in diagnostic reasoning: a randomised controlled experiment. BMJ Qual. Saf. 29, 550–559 (2020).
Mamede, S. et al. How can students’ diagnostic competence benefit most from practice with clinical cases? The effects of structured reflection on future diagnosis of the same and novel diseases. Acad. Med. 89, 121–127 (2014).
Mamede, S., Schmidt, H. G. & Penaforte, J. C. Effects of reflective practice on the accuracy of medical diagnoses. Med. Educ. 42, 468–475 (2008).
Norman, G. R. et al. The causes of errors in clinical reasoning: cognitive biases, knowledge deficits, and dual process thinking. Acad. Med. 92, 23–30 (2017).
Shao, Z. et al. DeepSeekMath: pushing the limits of mathematical reasoning in open language models. Preprint at https://arxiv.org/abs/2402.03300 (2024).
Singhal, K. et al. Large language models encode clinical knowledge. Nature 620, 172–180 (2023).
Yao, S. et al. ReAct: synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR, 2023).
Bakken, S. AI in health: keeping the human in the loop. J. Am. Med. Inform. Assoc. 30, 1225–1226 (2023).
For trustworthy AI, keep the human in the loop. Nat. Med. 31, 3207 (2025).
Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K. & Yao, S. Reflexion: language agents with verbal reinforcement learning. Adv. Neural Inf. Process. Syst. 36, 8634–8652 (2023).
Rodman, A. & Topol, E. J. Is generative artificial intelligence capable of clinical reasoning? Lancet 405, 689 (2025).
Zou, J. & Topol, E. J. The rise of agentic AI teammates in medicine. Lancet 405, 457 (2025).
Tordjman, M. et al. Comparative benchmarking of the DeepSeek large language model on medical tasks and clinical reasoning. Nat. Med. 31, 2550–2555 (2025).
Sandmann, S. et al. Benchmark evaluation of DeepSeek large language models in clinical decision-making. Nat. Med. 31, 2546–2549 (2025).
McDuff, D. et al. Towards accurate differential diagnosis with large language models. Nature 642, 451–457 (2025).
Tu, T. et al. Towards conversational diagnostic artificial intelligence. Nature 642, 442–450 (2025).
Johri, S. et al. An evaluation framework for clinical use of large language models in patient interaction tasks. Nat. Med. 31, 77–86 (2025).
Johnson, A. E. W. et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci. Data 10, 1 (2023).
Tomczak, K., Czerwińska, P. & Wiznerowicz, M. The Cancer Genome Atlas (TCGA): an immeasurable source of knowledge. Contemp. Oncol. 2015, 68–77 (2015).
Roberts, R. J. PubMed central: the GenBank of the published literature. Proc. Natl Acad. Sci. USA 98, 381–382 (2001).
Chakradhar, S. Predictable response: finding optimal drugs and doses using artificial intelligence. Nat. Med. 23, 1244–1247 (2017).
Lek, M. et al. Analysis of protein-coding genetic variation in 60,706 humans. Nature 536, 285–291 (2016).
Jiang, L. Y. et al. Health system-scale language models are all-purpose prediction engines. Nature 619, 357–362 (2023).
Asai, A., Wu, Z., Wang, Y., Sil, A. & Hajishirzi, H. Self-RAG: learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations (ICLR, 2023).
Kim, T. et al. MindfulDiary: harnessing large language model to support psychiatric patients’ journaling. In Proc. 2024 CHI Conference on Human Factors in Computing Systems 1–20 (ACM, 2024).
Ni, Y., Chen, Y., Ding, R. & Ni, S. Beatrice: a chatbot for collecting psychoecological data and providing QA capabilities. In Proc. 16th International Conference on Pervasive Technologies Related to Assistive Environments 429–435 (ACM, 2023).
Holderried, F. et al. A generative pretrained transformer (GPT)—powered chatbot as a simulated patient to practice history taking: prospective, mixed methods study. JMIR Med. Educ. 10, e53961 (2024).
Rodin, G. et al. Clinician–patient communication: a systematic review. Support. Care Cancer 17, 627–644 (2009).
Liu, S., McCoy, A. B. & Wright, A. Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines. J. Am. Med. Inform. Assoc. 32, 605–615 (2025).
Qiu, J. et al. LLM-based agentic systems in medicine and healthcare. Nat. Mach. Intell. 6, 1418–1420 (2024).
Ting, D. S. W. et al. Development and validation of a deep learning system for diabetic retinopathy and related eye diseases using retinal images from multiethnic populations with diabetes. JAMA 318, 2211–2223 (2017).
Liu, Y. et al. A deep learning system for differential diagnosis of skin diseases. Nat. Med. 26, 900–908 (2020).
Groh, M. et al. Deep learning-aided decision support for diagnosis of skin disease across skin tones. Nat. Med. 30, 573–583 (2024).
Lu, M. Y. et al. A multimodal generative AI copilot for human pathology. Nature 634, 466–473 (2024).
Tiu, E. et al. Expert-level detection of pathologies from unannotated chest X-ray images via self-supervised learning. Nat. Biomed. Eng. 6, 1399–1406 (2022).
Zhou, H.-Y. et al. A transformer-based representation-learning model with unified processing of multimodal input for clinical diagnostics. Nat. Biomed. Eng. 7, 743–755 (2023).
DeGrave, A. J., Cai, Z. R., Janizek, J. D., Daneshjou, R. & Lee, S.-I. Auditing the inference processes of medical-image classifiers by leveraging generative AI and the expertise of physicians. Nat. Biomed. Eng. 9, 294–306 (2025).
Brodeur, P. G. et al. Superhuman performance of a large language model on the reasoning tasks of a physician. Preprint at https://arxiv.org/abs/2412.10849 (2024).
Loh, H. W. et al. Application of explainable artificial intelligence for healthcare: a systematic review of the last decade (2011–2022). Comput. Methods Programs Biomed. 226, 107161 (2022).
Saraswat, D. et al. Explainable AI for healthcare 5.0: opportunities and challenges. IEEE Access 10, 84486–84517 (2022).
Schork, N. J. Artificial intelligence and personalized medicine. Precis. Med. Cancer Therapy 178, 265–283 (2019).
Parekh, A.-D. E., Shaikh, O. A., Simran, S., Manan, S. & Hasibuzzaman, M. A. Artificial intelligence (AI) in personalized medicine: AI-generated personalized therapy regimens based on genetic and medical history. Ann. Med. Surg. 85, 5831–5833 (2023).
Guk, K. et al. Evolution of wearable devices with real-time disease monitoring for personalized healthcare. Nanomaterials 9, 813 (2019).
Gao, S. et al. TxAgent: an AI agent for therapeutic reasoning across a universe of tools. Preprint at https://arxiv.org/abs/2503.10970 (2025).
Ji, C., Jiang, T., Liu, L., Zhang, J. & You, L. Continuous glucose monitoring combined with artificial intelligence: redefining the pathway for prediabetes management. Front. Endocrinol. 16, 1571362 (2025).
Subbiah, V. The next generation of evidence-based medicine. Nat. Med. 29, 49–58 (2023).
Wang, H. et al. Scientific discovery in the age of artificial intelligence. Nature 620, 47–60 (2023).
Gao, S. et al. Empowering biomedical discovery with AI agents. Cell 187, 6125–6151 (2024).
Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021).
Baek, M. et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science 373, 871–876 (2021).
Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100 (2023).
Kortemme, T. De novo protein design—from new structures to programmable functions. Cell 187, 526–544 (2024).
Swanson, K., Wu, W., Bulaong, N. L., Pak, J. E. & Zou, J. The virtual lab of AI agents designs new SARS-COV-2 nanobodies. Nature 646, 716–723 (2025).
Vaswani, A. et al. Attention is all you need. Adv. Neural Inf. Process. Syst. 30, 6000–6010 (2017).
Yu, Q. et al. DAPO: an open-source LLM reinforcement learning system at scale. Adv. Neural Inf. Process. Syst. 38, 113222–113244 (2026).
Schulman, J., Wolski, F., Dhariwal, P., Radford, A. & Klimov, O. Proximal policy optimization algorithms. Preprint at https://arxiv.org/abs/1707.06347 (2017).
Muennighoff, N. et al. s1: simple test-time scaling. In Proc. 2025 Conference on Empirical Methods in Natural Language Processing 20286–20332 (ACL, 2025).
Huang, X., Wu, J., Liu, H., Tang, X. & Zhou, Y. m1: unleash the potential of test-time scaling for medical reasoning with large language models. In Proc. Machine Learning for Health Vol. 297, 369–383 (PMLR, 2025).
Liévin, V., Hother, C. E., Motzfeldt, A. G. & Winther, O. Can large language models reason about medical questions? Patterns 5, 100943 (2024).
Nori, H. et al. Can generalist foundation models outcompete special-purpose tuning? Case study in medicine. Preprint at https://arxiv.org/abs/2311.16452 (2023).
Sonoda, Y. et al. Structured clinical reasoning prompt enhances LLM’s diagnostic capabilities in diagnosis please quiz cases. Jpn J. Radiol. 43, 586–592 (2025).
Savage, T., Nayak, A., Gallo, R., Rangan, E. & Chen, J. H. Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine. npj Digit. Med. 7, 20 (2024).
Yuksekgonul, M. et al. Optimizing generative ai by backpropagating language model feedback. Nature 639, 609–616 (2025).
Aali, A. et al. Prompt optimization improves robustness of language model benchmarks for medical tasks. In Machine Learning for Health (2025).
Bogireddy, S. P. T. R. et al. Neural at ArchEHR-QA 2025: agentic prompt optimization for evidence-grounded clinical question answering. In Proc. 24th Workshop on Biomedical Language Processing (Shared Tasks) 104–109 (ACL, 2025).
Khattab, O. et al. DSPy: compiling declarative language model calls into state-of-the-art pipelines. In The Twelfth International Conference on Learning Representations (ICLR, 2024).
Kim, Y. et al. MDAgents: an adaptive collaboration of llms for medical decision-making. Adv. Neural Inf. Process. Syst. 37, 79410–79452 (2024).
Li, X., Zou, H. & Liu, P. ToRL: scaling tool-integrated RL. Preprint at https://arxiv.org/abs/2503.23383 (2025).
Jin, B. et al. Search-R1: training LLMs to reason and leverage search engines with reinforcement learning. In Second Conference on Language Modeling (2025).
Chen, M. et al. Learning to reason with search for LLMs via reinforcement learning. Adv. Neural Inf. Process. Syst. 38, 85287–85307 (2026).
Wang, H. et al. OTC: optimal tool calls via reinforcement learning. Preprint at https://arxiv.org/abs/2504.05349 (2025).
Zheng, Q. et al. End-to-end agentic RAG system training for traceable diagnostic reasoning. Preprint at https://arxiv.org/abs/2508.15746 (2025).
Gulshan, V. et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA 316, 2402–2410 (2016).
Gilson, A. et al. How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med. Educ. 9, e45312 (2023).
Liu, N., Zhang, Z., Ho, A. F. W. & Ong, M. E. H. Artificial intelligence in emergency medicine. J. Emerg. Crit. Care Med. 2, 82 (2018).
De novo classification request. FDA https://www.fda.gov/medical-devices/premarket-submissions-selecting-and-preparing-correct-submission/de-novo-classification-request (2020).
McNamara, S. L., Yi, P. H. & Lotter, W. The clinician–AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation. npj Digit. Med. 7, 80 (2024).
Feng, J. et al. Clinical artificial intelligence quality improvement: towards continual monitoring and updating of AI algorithms in healthcare. npj Digit. Med. 5, 66 (2022).
Tai-Seale, M. et al. AI-generated draft replies integrated into health records and physicians’ electronic communication. JAMA Netw. Open 7, e246565–e246565 (2024).
Yin, J., Ngiam, K. Y., Tan, S. S.-L. & Teo, H. H. Designing AI-based work processes: how the timing of AI advice affects diagnostic decision making. Manag. Sci. 71, 8995–9868 (2025).
Vaccaro, M., Almaatouq, A. & Malone, T. When combinations of humans and AI are useful: a systematic review and meta-analysis. Nat. Hum. Behav. 8, 2293–2303 (2024).
Turpin, M., Michael, J., Perez, E. & Bowman, S. Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting. Adv. Neural Inf. Process. Syst. 36, 74952–74965 (2023).
Perrier, E. Typed chain-of-thought: a Curry–Howard framework for verifying LLM reasoning. Preprint at https://arxiv.org/abs/2510.01069 (2025).
Lee, J. & Hockenmaier, J. Evaluating step-by-step reasoning traces: a survey. In Findings of the Association for Computational Linguistics: EMNLP 1789–1814 (ACL, 2025).
Ling, Z. et al. Deductive verification of chain-of-thought reasoning. Adv. Neural Inf. Process. Syst. 36, 36407–36433 (2023).
Sag, M. Copyright safety for generative AI. Houst. Law Rev. 61, 295 (2023).
Giuffrè, M. & Shung, D. L. Harnessing the power of synthetic data in healthcare: innovation, application, and privacy. npj Digit. Med. 6, 186 (2023).
Seyyed-Kalantari, L., Zhang, H., McDermott, M. B., Chen, I. Y. & Ghassemi, M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat. Med. 27, 2176–2182 (2021).
Wongvibulsin, S. et al. Current state of dermatology mobile applications with artificial intelligence features. JAMA Dermatol. 160, 646–650 (2024).
Tanno, R. et al. Collaboration between clinicians and vision–language models in radiology report generation. Nat. Med. 31, 599–608 (2025).
Bharadwaj, P. et al. Unlocking the value: quantifying the return on investment of hospital artificial intelligence. J. Am. Coll. Radiol. 21, 1677–1685 (2024).
Reardon, S. Rise of robot radiologists. Nature 576, S54–S58 (2019).
Robert, D. et al. Effect of artificial intelligence as a second reader on the lung nodule detection and localization accuracy of radiologists and non-radiology physicians in chest radiographs: a multicenter reader study. Acad. Radiol. 32, 1706–1717 (2025).