Jaech, A. et al. OpenAI o1 system card. Preprint at https://arxiv.org/abs/2412.16720 (2024).

Guo, D. et al. DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature 645, 633–638 (2025).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Trinh, T. H., Wu, Y., Le, Q. V., He, H. & Luong, T. Solving olympiad geometry without human demonstrations. Nature 625, 476–482 (2024).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Rajpurkar, P., Chen, E., Banerjee, O. & Topol, E. J. AI in health and medicine. Nat. Med. 28, 31–38 (2022).

Article 
CAS 
PubMed 

Google Scholar
 

Mamede, S. et al. Immunising’physicians against availability bias in diagnostic reasoning: a randomised controlled experiment. BMJ Qual. Saf. 29, 550–559 (2020).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Mamede, S. et al. How can students’ diagnostic competence benefit most from practice with clinical cases? The effects of structured reflection on future diagnosis of the same and novel diseases. Acad. Med. 89, 121–127 (2014).

Article 
PubMed 

Google Scholar
 

Mamede, S., Schmidt, H. G. & Penaforte, J. C. Effects of reflective practice on the accuracy of medical diagnoses. Med. Educ. 42, 468–475 (2008).

Article 
PubMed 

Google Scholar
 

Norman, G. R. et al. The causes of errors in clinical reasoning: cognitive biases, knowledge deficits, and dual process thinking. Acad. Med. 92, 23–30 (2017).

Article 
PubMed 

Google Scholar
 

Shao, Z. et al. DeepSeekMath: pushing the limits of mathematical reasoning in open language models. Preprint at https://arxiv.org/abs/2402.03300 (2024).

Singhal, K. et al. Large language models encode clinical knowledge. Nature 620, 172–180 (2023).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Yao, S. et al. ReAct: synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR, 2023).

Bakken, S. AI in health: keeping the human in the loop. J. Am. Med. Inform. Assoc. 30, 1225–1226 (2023).

Article 
PubMed 
PubMed Central 

Google Scholar
 

For trustworthy AI, keep the human in the loop. Nat. Med. 31, 3207 (2025).

Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K. & Yao, S. Reflexion: language agents with verbal reinforcement learning. Adv. Neural Inf. Process. Syst. 36, 8634–8652 (2023).


Google Scholar
 

Rodman, A. & Topol, E. J. Is generative artificial intelligence capable of clinical reasoning? Lancet 405, 689 (2025).

Article 
PubMed 

Google Scholar
 

Zou, J. & Topol, E. J. The rise of agentic AI teammates in medicine. Lancet 405, 457 (2025).

Article 
PubMed 

Google Scholar
 

Tordjman, M. et al. Comparative benchmarking of the DeepSeek large language model on medical tasks and clinical reasoning. Nat. Med. 31, 2550–2555 (2025).

Article 
CAS 
PubMed 

Google Scholar
 

Sandmann, S. et al. Benchmark evaluation of DeepSeek large language models in clinical decision-making. Nat. Med. 31, 2546–2549 (2025).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

McDuff, D. et al. Towards accurate differential diagnosis with large language models. Nature 642, 451–457 (2025).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Tu, T. et al. Towards conversational diagnostic artificial intelligence. Nature 642, 442–450 (2025).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Johri, S. et al. An evaluation framework for clinical use of large language models in patient interaction tasks. Nat. Med. 31, 77–86 (2025).

Article 
CAS 
PubMed 

Google Scholar
 

Johnson, A. E. W. et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci. Data 10, 1 (2023).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Tomczak, K., Czerwińska, P. & Wiznerowicz, M. The Cancer Genome Atlas (TCGA): an immeasurable source of knowledge. Contemp. Oncol. 2015, 68–77 (2015).


Google Scholar
 

Roberts, R. J. PubMed central: the GenBank of the published literature. Proc. Natl Acad. Sci. USA 98, 381–382 (2001).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Chakradhar, S. Predictable response: finding optimal drugs and doses using artificial intelligence. Nat. Med. 23, 1244–1247 (2017).

Article 
CAS 
PubMed 

Google Scholar
 

Lek, M. et al. Analysis of protein-coding genetic variation in 60,706 humans. Nature 536, 285–291 (2016).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Jiang, L. Y. et al. Health system-scale language models are all-purpose prediction engines. Nature 619, 357–362 (2023).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Asai, A., Wu, Z., Wang, Y., Sil, A. & Hajishirzi, H. Self-RAG: learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations (ICLR, 2023).

Kim, T. et al. MindfulDiary: harnessing large language model to support psychiatric patients’ journaling. In Proc. 2024 CHI Conference on Human Factors in Computing Systems 1–20 (ACM, 2024).

Ni, Y., Chen, Y., Ding, R. & Ni, S. Beatrice: a chatbot for collecting psychoecological data and providing QA capabilities. In Proc. 16th International Conference on Pervasive Technologies Related to Assistive Environments 429–435 (ACM, 2023).

Holderried, F. et al. A generative pretrained transformer (GPT)—powered chatbot as a simulated patient to practice history taking: prospective, mixed methods study. JMIR Med. Educ. 10, e53961 (2024).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Rodin, G. et al. Clinician–patient communication: a systematic review. Support. Care Cancer 17, 627–644 (2009).

PubMed 

Google Scholar
 

Liu, S., McCoy, A. B. & Wright, A. Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines. J. Am. Med. Inform. Assoc. 32, 605–615 (2025).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Qiu, J. et al. LLM-based agentic systems in medicine and healthcare. Nat. Mach. Intell. 6, 1418–1420 (2024).

Article 

Google Scholar
 

Ting, D. S. W. et al. Development and validation of a deep learning system for diabetic retinopathy and related eye diseases using retinal images from multiethnic populations with diabetes. JAMA 318, 2211–2223 (2017).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Liu, Y. et al. A deep learning system for differential diagnosis of skin diseases. Nat. Med. 26, 900–908 (2020).

Article 
CAS 
PubMed 

Google Scholar
 

Groh, M. et al. Deep learning-aided decision support for diagnosis of skin disease across skin tones. Nat. Med. 30, 573–583 (2024).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Lu, M. Y. et al. A multimodal generative AI copilot for human pathology. Nature 634, 466–473 (2024).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Tiu, E. et al. Expert-level detection of pathologies from unannotated chest X-ray images via self-supervised learning. Nat. Biomed. Eng. 6, 1399–1406 (2022).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Zhou, H.-Y. et al. A transformer-based representation-learning model with unified processing of multimodal input for clinical diagnostics. Nat. Biomed. Eng. 7, 743–755 (2023).

Article 
PubMed 

Google Scholar
 

DeGrave, A. J., Cai, Z. R., Janizek, J. D., Daneshjou, R. & Lee, S.-I. Auditing the inference processes of medical-image classifiers by leveraging generative AI and the expertise of physicians. Nat. Biomed. Eng. 9, 294–306 (2025).

Article 
PubMed 

Google Scholar
 

Brodeur, P. G. et al. Superhuman performance of a large language model on the reasoning tasks of a physician. Preprint at https://arxiv.org/abs/2412.10849 (2024).

Loh, H. W. et al. Application of explainable artificial intelligence for healthcare: a systematic review of the last decade (2011–2022). Comput. Methods Programs Biomed. 226, 107161 (2022).

Article 
PubMed 

Google Scholar
 

Saraswat, D. et al. Explainable AI for healthcare 5.0: opportunities and challenges. IEEE Access 10, 84486–84517 (2022).

Article 

Google Scholar
 

Schork, N. J. Artificial intelligence and personalized medicine. Precis. Med. Cancer Therapy 178, 265–283 (2019).

Article 
CAS 

Google Scholar
 

Parekh, A.-D. E., Shaikh, O. A., Simran, S., Manan, S. & Hasibuzzaman, M. A. Artificial intelligence (AI) in personalized medicine: AI-generated personalized therapy regimens based on genetic and medical history. Ann. Med. Surg. 85, 5831–5833 (2023).

Article 

Google Scholar
 

Guk, K. et al. Evolution of wearable devices with real-time disease monitoring for personalized healthcare. Nanomaterials 9, 813 (2019).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Gao, S. et al. TxAgent: an AI agent for therapeutic reasoning across a universe of tools. Preprint at https://arxiv.org/abs/2503.10970 (2025).

Ji, C., Jiang, T., Liu, L., Zhang, J. & You, L. Continuous glucose monitoring combined with artificial intelligence: redefining the pathway for prediabetes management. Front. Endocrinol. 16, 1571362 (2025).

Article 

Google Scholar
 

Subbiah, V. The next generation of evidence-based medicine. Nat. Med. 29, 49–58 (2023).

Article 
CAS 
PubMed 

Google Scholar
 

Wang, H. et al. Scientific discovery in the age of artificial intelligence. Nature 620, 47–60 (2023).

Article 
CAS 
PubMed 

Google Scholar
 

Gao, S. et al. Empowering biomedical discovery with AI agents. Cell 187, 6125–6151 (2024).

Article 
CAS 
PubMed 

Google Scholar
 

Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Baek, M. et al. Accurate prediction of protein structures and interactions using a three-track neural network. Science 373, 871–876 (2021).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Watson, J. L. et al. De novo design of protein structure and function with RFdiffusion. Nature 620, 1089–1100 (2023).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Kortemme, T. De novo protein design—from new structures to programmable functions. Cell 187, 526–544 (2024).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Swanson, K., Wu, W., Bulaong, N. L., Pak, J. E. & Zou, J. The virtual lab of AI agents designs new SARS-COV-2 nanobodies. Nature 646, 716–723 (2025).

Article 
CAS 
PubMed 

Google Scholar
 

Vaswani, A. et al. Attention is all you need. Adv. Neural Inf. Process. Syst. 30, 6000–6010 (2017).


Google Scholar
 

Yu, Q. et al. DAPO: an open-source LLM reinforcement learning system at scale. Adv. Neural Inf. Process. Syst. 38, 113222–113244 (2026).


Google Scholar
 

Schulman, J., Wolski, F., Dhariwal, P., Radford, A. & Klimov, O. Proximal policy optimization algorithms. Preprint at https://arxiv.org/abs/1707.06347 (2017).

Muennighoff, N. et al. s1: simple test-time scaling. In Proc. 2025 Conference on Empirical Methods in Natural Language Processing 20286–20332 (ACL, 2025).

Huang, X., Wu, J., Liu, H., Tang, X. & Zhou, Y. m1: unleash the potential of test-time scaling for medical reasoning with large language models. In Proc. Machine Learning for Health Vol. 297, 369–383 (PMLR, 2025).

Liévin, V., Hother, C. E., Motzfeldt, A. G. & Winther, O. Can large language models reason about medical questions? Patterns 5, 100943 (2024).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Nori, H. et al. Can generalist foundation models outcompete special-purpose tuning? Case study in medicine. Preprint at https://arxiv.org/abs/2311.16452 (2023).

Sonoda, Y. et al. Structured clinical reasoning prompt enhances LLM’s diagnostic capabilities in diagnosis please quiz cases. Jpn J. Radiol. 43, 586–592 (2025).

PubMed 

Google Scholar
 

Savage, T., Nayak, A., Gallo, R., Rangan, E. & Chen, J. H. Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine. npj Digit. Med. 7, 20 (2024).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Yuksekgonul, M. et al. Optimizing generative ai by backpropagating language model feedback. Nature 639, 609–616 (2025).

Article 
CAS 
PubMed 

Google Scholar
 

Aali, A. et al. Prompt optimization improves robustness of language model benchmarks for medical tasks. In Machine Learning for Health (2025).

Bogireddy, S. P. T. R. et al. Neural at ArchEHR-QA 2025: agentic prompt optimization for evidence-grounded clinical question answering. In Proc. 24th Workshop on Biomedical Language Processing (Shared Tasks) 104–109 (ACL, 2025).

Khattab, O. et al. DSPy: compiling declarative language model calls into state-of-the-art pipelines. In The Twelfth International Conference on Learning Representations (ICLR, 2024).

Kim, Y. et al. MDAgents: an adaptive collaboration of llms for medical decision-making. Adv. Neural Inf. Process. Syst. 37, 79410–79452 (2024).


Google Scholar
 

Li, X., Zou, H. & Liu, P. ToRL: scaling tool-integrated RL. Preprint at https://arxiv.org/abs/2503.23383 (2025).

Jin, B. et al. Search-R1: training LLMs to reason and leverage search engines with reinforcement learning. In Second Conference on Language Modeling (2025).

Chen, M. et al. Learning to reason with search for LLMs via reinforcement learning. Adv. Neural Inf. Process. Syst. 38, 85287–85307 (2026).


Google Scholar
 

Wang, H. et al. OTC: optimal tool calls via reinforcement learning. Preprint at https://arxiv.org/abs/2504.05349 (2025).

Zheng, Q. et al. End-to-end agentic RAG system training for traceable diagnostic reasoning. Preprint at https://arxiv.org/abs/2508.15746 (2025).

Gulshan, V. et al. Development and validation of a deep learning algorithm for detection of diabetic retinopathy in retinal fundus photographs. JAMA 316, 2402–2410 (2016).

Article 
PubMed 

Google Scholar
 

Gilson, A. et al. How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)? The implications of large language models for medical education and knowledge assessment. JMIR Med. Educ. 9, e45312 (2023).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Liu, N., Zhang, Z., Ho, A. F. W. & Ong, M. E. H. Artificial intelligence in emergency medicine. J. Emerg. Crit. Care Med. 2, 82 (2018).

Article 

Google Scholar
 

De novo classification request. FDA https://www.fda.gov/medical-devices/premarket-submissions-selecting-and-preparing-correct-submission/de-novo-classification-request (2020).

McNamara, S. L., Yi, P. H. & Lotter, W. The clinician–AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation. npj Digit. Med. 7, 80 (2024).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Feng, J. et al. Clinical artificial intelligence quality improvement: towards continual monitoring and updating of AI algorithms in healthcare. npj Digit. Med. 5, 66 (2022).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Tai-Seale, M. et al. AI-generated draft replies integrated into health records and physicians’ electronic communication. JAMA Netw. Open 7, e246565–e246565 (2024).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Yin, J., Ngiam, K. Y., Tan, S. S.-L. & Teo, H. H. Designing AI-based work processes: how the timing of AI advice affects diagnostic decision making. Manag. Sci. 71, 8995–9868 (2025).


Google Scholar
 

Vaccaro, M., Almaatouq, A. & Malone, T. When combinations of humans and AI are useful: a systematic review and meta-analysis. Nat. Hum. Behav. 8, 2293–2303 (2024).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Turpin, M., Michael, J., Perez, E. & Bowman, S. Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting. Adv. Neural Inf. Process. Syst. 36, 74952–74965 (2023).


Google Scholar
 

Perrier, E. Typed chain-of-thought: a Curry–Howard framework for verifying LLM reasoning. Preprint at https://arxiv.org/abs/2510.01069 (2025).

Lee, J. & Hockenmaier, J. Evaluating step-by-step reasoning traces: a survey. In Findings of the Association for Computational Linguistics: EMNLP 1789–1814 (ACL, 2025).

Ling, Z. et al. Deductive verification of chain-of-thought reasoning. Adv. Neural Inf. Process. Syst. 36, 36407–36433 (2023).


Google Scholar
 

Sag, M. Copyright safety for generative AI. Houst. Law Rev. 61, 295 (2023).


Google Scholar
 

Giuffrè, M. & Shung, D. L. Harnessing the power of synthetic data in healthcare: innovation, application, and privacy. npj Digit. Med. 6, 186 (2023).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Seyyed-Kalantari, L., Zhang, H., McDermott, M. B., Chen, I. Y. & Ghassemi, M. Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nat. Med. 27, 2176–2182 (2021).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Wongvibulsin, S. et al. Current state of dermatology mobile applications with artificial intelligence features. JAMA Dermatol. 160, 646–650 (2024).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Tanno, R. et al. Collaboration between clinicians and vision–language models in radiology report generation. Nat. Med. 31, 599–608 (2025).

Article 
CAS 
PubMed 

Google Scholar
 

Bharadwaj, P. et al. Unlocking the value: quantifying the return on investment of hospital artificial intelligence. J. Am. Coll. Radiol. 21, 1677–1685 (2024).

Article 
PubMed 

Google Scholar
 

Reardon, S. Rise of robot radiologists. Nature 576, S54–S58 (2019).

Article 
CAS 
PubMed 

Google Scholar
 

Robert, D. et al. Effect of artificial intelligence as a second reader on the lung nodule detection and localization accuracy of radiologists and non-radiology physicians in chest radiographs: a multicenter reader study. Acad. Radiol. 32, 1706–1717 (2025).

Article 
PubMed 

Google Scholar