{"id":425911,"date":"2026-04-30T22:17:07","date_gmt":"2026-04-30T22:17:07","guid":{"rendered":"https:\/\/www.newsbeep.com\/ie\/425911\/"},"modified":"2026-04-30T22:17:07","modified_gmt":"2026-04-30T22:17:07","slug":"ai-just-beat-doctors-at-diagnosing-er-patients-dont-get-all-excited","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ie\/425911\/","title":{"rendered":"AI Just Beat Doctors at Diagnosing ER Patients. Don&#8217;t Get All Excited"},"content":{"rendered":"<p>Emergency departments and other clinical settings across the world are now one step closer to sounding like the cockpit of the Millennium Falcon\u2014with human doctors soliciting advice from, bickering with, and not infrequently trusting the guidance of their opinionated AI colleagues.<\/p>\n<p>Researchers at Harvard and Boston\u2019s Beth Israel Deaconess Medical Center have successfully tested an advanced large language model (LLM) AI against two attending physicians (humans) in their performance diagnosing incoming emergency room patients at the triage phase.<\/p>\n<p>The LLM, <a href=\"https:\/\/gizmodo.com\/why-does-chatgpts-algorithm-think-in-chinese-2000550311#:~:text=OpenAI%20recently%20unveiled,several%20other%20languages.\" rel=\"nofollow noopener\" target=\"_blank\">OpenAI\u2019s first so-called \u201creasoning\u201d model o1-preview<\/a>, made the correct call in 67.1% of the 76 actual emergency department cases put to it, with what the researchers called \u201cexact or a very close\u201d diagnostic accuracy in the new <a href=\"http:\/\/www.science.org\/doi\/10.1126\/science.adz4433\" rel=\"nofollow noopener\" target=\"_blank\">study<\/a>, published today in the journal Science. Two expert physicians sourced from elite university medical institutions, however, only scored 55.3% and 50.0% accuracy, respectively, with blinded physician reviewers unable to tell these o1 and human-made diagnoses apart.<\/p>\n<p>The new study also pitted o1 and OpenAI\u2019s prior non-reasoning LLMs, like ChatGPT-4, against physicians\u2019 past testing baselines diagnosing 143 complex cases published as clinical vignettes in The New England Journal of Medicine.<\/p>\n<p>\u201co1-preview included the correct diagnosis in its differential in 78.3% of these cases,\u201d according to one of the study\u2019s lead authors, doctoral candidate Thomas Buckley with Harvard Medical School\u2019s Department of Biomedical Informatics, who spoke at a press briefing Tuesday.<\/p>\n<p>\u201cAnd when expanding to a differential diagnosis that would have been helpful,\u201d Buckley continued, \u201cwe found that o1-preview suggested a helpful diagnosis in 97.9% of cases.\u201d The results, he noted, not only outperformed ChatGPT-4 but also vastly outpaced a human physician baseline <a href=\"https:\/\/doi.org\/10.1038\/s41586-025-08869-4\" rel=\"nofollow noopener\" target=\"_blank\">published<\/a> in Nature, where physicians with the freedom to consult search engines and standard medical resources had an accuracy of 44.5%. (Although, this study included a larger and perhaps more thorny set of 302 clinical vignettes.)<\/p>\n<p> I, Robot, M.D. <\/p>\n<p>\u201cI don\u2019t think our findings mean that AI replaces doctors,\u201d study coauthor Arjun Manrai, who teaches biomedical informatics at Harvard, took pains to emphasize at the press briefing, \u201cdespite what some companies are likely to say.\u201d<\/p>\n<p>Manrai did, however, describe the team\u2019s results as evidence of a \u201creally profound change in technology that will reshape medicine,\u201d one that would require rigorous testing to verify their utility in actually making patient outcomes better.<\/p>\n<p>Two independent medical researchers, who <a href=\"http:\/\/www.science.org\/doi\/10.1126\/science.aeg8766\" rel=\"nofollow noopener\" target=\"_blank\">commented<\/a> on the new study in a piece published concurrently in Science, echoed this view. \u201cThe prevailing proposal for AI in health care is not replacement but collaboration,\u201d they noted, \u201cwith clinicians providing oversight, contextual judgment, and accountability.\u201d<\/p>\n<p>Study coauthor Adam Rodman, an internal medicine physician at Beth Israel, likened the possible legal status of AI diagnoses to the current paradigm with clinical decision support (CDS), already existing digital tools doctors use while retaining personal culpability for those choices.<\/p>\n<p>\u201cI will tell you, as a practicing physician, that would be a limitation to widespread adoption of all of this, if the regulatory system is \u2018Just trust me,\u2019\u201d Rodman said at the briefing. \u201cI would have to see extraordinarily strong evidence, such as a randomized controlled trial, where I would do that for my patients.\u201d<\/p>\n<p> Playing doctor <\/p>\n<p>Reasoning models, like o1-preview, differ from the AI chatbots you might be used to in that these LLMs have been built to work through problems in structured steps, mirroring more deductive thinking, before delivering answers to a prompt. The system still has its limitations, which, according to the researchers, include real difficulty diagnosing medical cases involving multimodal input, meaning images and audio evidence that would easily help a human doctor diagnose a patient\u2019s case.<\/p>\n<p>\u201cThey\u2019re underperforming on most medical imaging benchmarks,\u201d Buckley said. \u201cI think a really active area of research over the next decade is how do we improve the multimodal integration capabilities of these models.\u201d<\/p>\n<p>Yujin Potter\u2014an AI research scientist at the University of California, Berkeley, who reviewed the new study for Gizmodo\u2014noted that the team\u2019s finished paper was quiet on more troubling issues now known to plague AI. Potter, who\u2019s not involved with the new research,\u00a0co-published a study in March <a href=\"https:\/\/rdi.berkeley.edu\/blog\/peer-preservation\/\" rel=\"nofollow noopener\" target=\"_blank\">detailing<\/a> how teams of AI can spontaneously develop and act on their own goals when tasked to work in coordination, <a href=\"https:\/\/gizmodo.com\/llms-will-protect-each-other-if-threatened-study-finds-2000741634\" rel=\"nofollow noopener\" target=\"_blank\">actively deceiving their human users and exfiltrating files to hide on different servers<\/a>.<\/p>\n<p>\u201cThis paper is informative. It\u2019s good. But also, this actually means that we also need to understand AI safety better,\u201d Potter told Gizmodo. \u201cPeople should keep in their mind that AI can also hallucinate and give them the wrong information\u2014and even malicious or misaligned AI can manipulate them.\u201d<\/p>\n<p>At the Tuesday briefing, Buckley acknowledged that he and his colleagues \u201cdidn\u2019t formally measure the hallucination rate of these models.\u201d<\/p>\n<p>\u201cWe do know that models such as o1 do hallucinate,\u201d Buckley added, \u201cbut in the significant majority of cases, we are finding that the model is suggesting something at least helpful, and then in a huge amount of cases, it\u2019s suggesting the exact diagnosis in the original case.\u201d<\/p>\n<p>Manrai, Buckley\u2019s coauthor, added: \u201cMy mantra is still \u2018trust, but verify.\u2019\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"Emergency departments and other clinical settings across the world are now one step closer to sounding like the&hellip;\n","protected":false},"author":2,"featured_media":425912,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[34],"tags":[218,103,397,396,61,60,186171,1682,186172],"class_list":["post-425911","post","type-post","status-publish","format-standard","has-post-thumbnail","category-healthcare","tag-artificial-intelligence","tag-health","tag-health-care","tag-healthcare","tag-ie","tag-ireland","tag-medical-innovations","tag-openai","tag-reasoning-model-llms"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/425911","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/comments?post=425911"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/425911\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media\/425912"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media?parent=425911"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/categories?post=425911"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/tags?post=425911"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}