{"id":653500,"date":"2026-05-06T14:44:38","date_gmt":"2026-05-06T14:44:38","guid":{"rendered":"https:\/\/www.newsbeep.com\/au\/653500\/"},"modified":"2026-05-06T14:44:38","modified_gmt":"2026-05-06T14:44:38","slug":"ai-outperforms-doctors-in-diagnosis-tests-but-dont-expect-it-to-replace-your-doctor","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/au\/653500\/","title":{"rendered":"AI outperforms doctors in diagnosis tests \u2014 but don&#8217;t expect it to replace your doctor"},"content":{"rendered":"<p>A new study that examined how well artificial intelligence performed in an emergency room setting found that it outperformed doctors at diagnosing patients.\u00a0<\/p>\n<p>The study, <a href=\"https:\/\/www.science.org\/doi\/10.1126\/science.adz4433#sec-1\" rel=\"nofollow noopener\" target=\"_blank\">published in Science<\/a> on Thursday, evaluated OpenAI\u2019s o1 model, which the company released in 2024. The model is a reasoning-focused AI specifically designed to excel at complex, structured problems. This makes it fairly different from chatbots like ChatGPT, which OpenAI designed as a generalist.<\/p>\n<p>Despite the positive results from the study, the researchers emphasized the study\u2019s limitations and raised concerns that their results may be used by others to suggest AI replace doctors. Dr. Adam Rodman, a general internis\u00ad\u00ad\u00adt and medical educator at Beth Israel Deaconess Medical Center and the co-author of the study, <a href=\"http:\/\/vox.com\/health\/487425\/open-ai-chatgpt-diagnosis-symptoms-second-opinion-study\" rel=\"nofollow noopener\" target=\"_blank\">said<\/a> he gets \u201ca little bit queasy about how some of these results might be used.\u201d<\/p>\n<p>\t\t<img loading=\"lazy\" decoding=\"async\" class=\"wp-block-san-app-download__qr\" src=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2026\/04\/app-download-block-qr-code.png\" alt=\"QR code for SAN app download\" width=\"80\" height=\"80\"\/><\/p>\n<p class=\"wp-block-san-app-download__title\">\n\t\t\tDownload the SAN app today to stay up-to-date with Unbiased. Straight Facts\u2122.\t\t<\/p>\n<p class=\"wp-block-san-app-download__subtitle\">\n\t\t\tPoint phone camera here\t\t<\/p>\n<p>What did the studies find?<\/p>\n<p>The research team tested the model in different ways. The first way tested against curated medical training cases. These cases are specifically designed to test doctors\u2019 diagnostic thinking and are just like the exams physicians take in school.\u00a0<\/p>\n<p>In these tests, the AI model consistently outperformed across these scenarios. But the way that test where the AI really excelled was with historical real-world ER cases.\u00a0<\/p>\n<p>During this test, researchers pulled patient cases from Beth Israel Deaconess Medical Center to get as close to real conditions as possible. The researchers noted that they gave the AI raw electronic health record data, which they described as \u201cmessy\u201d and similar to what actual doctors encounter.\u00a0<\/p>\n<p>The team tested the data at two different points: when the patient arrived at triage and then later when the patient was ready for admission. In the first test, the AI got the correct diagnosis 67% of the time, compared to 50% and 55% for the two human doctors it measured against. By the time the patient was ready for admission, the AI\u2019s diagnosis jumped to 81%, compared to 70% and 79% for the human doctors.\u00a0<\/p>\n<p>Although the tests were not conducted during the hospital visit, the researchers found that the AI was effective at making diagnoses.\u00a0<\/p>\n<p>\u201cWe can definitively say \u2026 reasoning models can meet that criteria for making diagnostic reasoning at the highest levels of human performance,\u201d Rodman said, according to Vox.\u00a0<\/p>\n<p>ChatGPT is still no doctor<\/p>\n<p>While specific models developed for the medical field performed well in tests, that doesn\u2019t mean ChatGPT or other large language models are a replacement for a general physician. Something the researchers specifically noted.\u00a0<\/p>\n<p>\u201dNo one should look at this and say we do not need doctors,\u201d Rodman said.<\/p>\n<p>A previous study published in February in <a href=\"https:\/\/www.nature.com\/articles\/s41591-026-04297-7\" rel=\"nofollow noopener\" target=\"_blank\">Nature Medicine<\/a> found that ChatGPT underestimated the severity of a patient\u2019s condition in 52% of cases.<\/p>\n<p>\t\t\tStart your day with fact-based news.<\/p>\n<p class=\"wp-block-san-san-inarticle-newsletter-signup__learn-more\">\n\t\t\t\t<a href=\"https:\/\/san.com\/newsletters\" target=\"_blank\" rel=\"nofollow noopener\">Learn more<\/a> about our emails. Unsubscribe Anytime.\n\t\t\t<\/p>\n<p>To test this, researchers gave the chatbot a range of medical scenarios, from non-urgent to medical emergencies, and evaluated its performance in assessment. In one example of the AI failing to notice the severity of the diagnosis, the bot told a patient on the verge of diabetic shock or respiratory failure to just monitor themselves instead of seeking immediate care. It also repeatedly failed to pick up on clear signs of suicidal ideation, a topic <a href=\"https:\/\/san.com\/cc\/ai-chatbots-are-too-agreeable-authorities-say-its-creating-deadly-outcomes\/#:~:text=Chatbots%20and%20suicides\" rel=\"nofollow noopener\" target=\"_blank\">Straight Arrow has previously reported on<\/a>.\u00a0<\/p>\n<p>The authors of the study published on Thursday have asked that their research be used in clinical trials testing how AI performs under real-world conditions before any hospital uses it to deploy more AI.\u00a0<\/p>\n<p>\u201cOur findings suggest the urgent need for prospective trials to evaluate these technologies in real-world patient care settings,\u201d the authors wrote.\u00a0<\/p>\n","protected":false},"excerpt":{"rendered":"A new study that examined how well artificial intelligence performed in an emergency room setting found that it&hellip;\n","protected":false},"author":2,"featured_media":653501,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[34],"tags":[64,63,137,500],"class_list":["post-653500","post","type-post","status-publish","format-standard","has-post-thumbnail","category-healthcare","tag-au","tag-australia","tag-health","tag-healthcare"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/653500","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/comments?post=653500"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/653500\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media\/653501"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media?parent=653500"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/categories?post=653500"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/tags?post=653500"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}