{"id":523966,"date":"2026-06-29T10:26:13","date_gmt":"2026-06-29T10:26:13","guid":{"rendered":"https:\/\/www.newsbeep.com\/ie\/523966\/"},"modified":"2026-06-29T10:26:13","modified_gmt":"2026-06-29T10:26:13","slug":"how-accurate-are-ai-chatbots-when-we-ask-them-a-medical-question-the-irish-times","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ie\/523966\/","title":{"rendered":"How accurate are AI chatbots when we ask them a medical question? \u2013 The Irish Times"},"content":{"rendered":"<p class=\"c-paragraph paywall \">When we use a search engine such as <a href=\"https:\/\/www.irishtimes.com\/tags\/google\/\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/www.irishtimes.com\/tags\/google\/\">Google<\/a> to look up information, it is now standard to see an artificial intelligence (<a href=\"https:\/\/www.irishtimes.com\/tags\/artificial-intelligence\/\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/www.irishtimes.com\/tags\/artificial-intelligence\/\">AI<\/a>) answer at the top of the results page. It is a measure of how much AI is part of our lives.<\/p>\n<p class=\"c-paragraph paywall \">But how accurate are AI chatbots when they are asked a <a href=\"https:\/\/www.irishtimes.com\/health\/\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/www.irishtimes.com\/health\/\">medical<\/a> question?<\/p>\n<p class=\"c-paragraph paywall \">A couple of recent research papers have come up with some disturbing answers. The authors of one study, <a href=\"https:\/\/bmjopen.bmj.com\/content\/16\/4\/e112695\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/bmjopen.bmj.com\/content\/16\/4\/e112695\">published in BMJ Open<\/a>, put five of the world\u2019s most popular chatbots through a systematic health-information stress test.<\/p>\n<p class=\"c-paragraph paywall \">The chatbots \u2013 ChatGPT, Gemini, Grok, Meta AI and DeepSeek \u2013 were each asked 50 health and medical questions covering topics such as cancer, vaccines, stem cells and nutrition. Two experts independently rated every answer. They found that nearly a fifth of the answers were highly problematic, half were problematic, and 30 per cent were somewhat problematic. None of the chatbots produced fully accurate reference lists. Chatbot performance varied by topic. They handled vaccines and cancer best \u2013 fields with large, well-structured bodies of research \u2013 yet still produced problematic answers roughly a quarter of the time.<\/p>\n<p class=\"c-paragraph paywall \">Significantly, the chatbots struggled most with open-ended questions \u2013 32 per cent of those answers were rated highly problematic, compared with just 7 per cent for closed questions. Most health queries people ask are open-ended. They tend not to ask chatbots true-or-false questions.<\/p>\n<p class=\"c-paragraph paywall \">When the researchers asked each chatbot for 10 scientific references, none managed a single fully accurate reference list in 25 attempts. A separate study published in Nature Medicine also found chatbot answers to be incomplete \u2013 but with an interesting twist. Dr Rebecca Payne, a GP and senior clinical lecturer at Bangor University in Wales, and colleagues gave participants brief descriptions of common medical situations. They were randomly assigned either to use one of three widely available chatbots or to rely on whatever sources they would normally use at home.<\/p>\n<p class=\"c-paragraph paywall \">After interacting with the chatbot, they were asked two questions: what condition might explain the symptoms? And where should they seek help?<\/p>\n<p class=\"c-paragraph paywall \">People who used chatbots were less likely to identify the correct condition than those who didn\u2019t. They were also no better at determining the right place to seek care than the control group. In other words, interacting with a chatbot did not help people make better health decisions. <\/p>\n<p class=\"c-paragraph paywall \">When the researchers then removed the human element and gave the same scenarios directly to the chatbots, their performance improved dramatically. Without human involvement, the models identified relevant conditions in the vast majority of cases and often suggested appropriate levels of care.<\/p>\n<p class=\"c-paragraph b-it-article-body__interstitial-link\">[\u00a0<a aria-label=\"Open related story\" class=\"c-link\" href=\"https:\/\/www.irishtimes.com\/opinion\/2025\/12\/15\/the-chatbot-will-see-you-now-is-this-the-future-of-irish-medicine\/\" rel=\"noreferrer nofollow noopener\" target=\"_blank\">The chatbot will see you now: is this the future of Irish medicine?Opens in new window<\/a>\u00a0]<\/p>\n<p class=\"c-paragraph paywall \">So why did the results deteriorate when people actually used the systems?<\/p>\n<p class=\"c-paragraph paywall \">Chatbots frequently mentioned the relevant diagnosis somewhere in the conversation, yet participants did not always notice or remember it when summarising their final answer. <\/p>\n<p class=\"c-paragraph paywall \">In other cases, users provided incomplete information or the chatbot misinterpreted key details. According to Dr Payne, the issue was not simply a failure of medical knowledge \u2013 it was a failure of communication between human and machine. <\/p>\n<p class=\"c-paragraph paywall \">Writing about her research <a href=\"https:\/\/theconversation.com\/why-ai-health-chatbots-wont-make-you-better-at-diagnosing-yourself-new-research-278049\" target=\"_blank\" rel=\"noreferrer nofollow noopener\" title=\"https:\/\/theconversation.com\/why-ai-health-chatbots-wont-make-you-better-at-diagnosing-yourself-new-research-278049\">in The Conversation<\/a>, Payne says the lesson from her study is not that AI has no place in healthcare. Rather, the key is understanding what these systems are currently good at and where their limitations lie.<\/p>\n<p class=\"c-paragraph b-it-article-body__interstitial-link\">[\u00a0<a aria-label=\"Open related story\" class=\"c-link\" href=\"https:\/\/www.irishtimes.com\/health\/2025\/12\/03\/doctors-reliance-on-ai-tools-could-erode-critical-thinking-experts-warn\/\" rel=\"noreferrer nofollow noopener\" target=\"_blank\">Doctors\u2019 reliance on AI tools could erode critical thinking, experts warnOpens in new window<\/a>\u00a0]<\/p>\n<p class=\"c-paragraph paywall \">\u201cOne useful way to think about today\u2019s chatbots is that they function more like secretaries than physicians. They are remarkably effective at organising information, summarising text and structuring complex documents. These are the kinds of tasks where language models are already proving useful within healthcare systems, for example in drafting clinical notes, summarising patient records or generating referral letters.\u201d<\/p>\n<p class=\"c-paragraph paywall \">The role of AI in medicine is likely to be more supportive than revolutionary in the near term. Chatbots should not be expected to act as the front door to healthcare. They are simply not ready to diagnose conditions or direct patients to the right level of care. <\/p>\n<p class=\"c-paragraph paywall \"><a href=\"https:\/\/www.irishtimes.com\/health\/your-wellness\/2026\/06\/29\/how-accurate-are-ai-chatbots-when-we-ask-them-a-medical-question\/mailto:mhouston@irishtimes.com\" rel=\"nofollow noopener\" title=\"https:\/\/www.irishtimes.com\/health\/your-wellness\/2026\/06\/29\/how-accurate-are-ai-chatbots-when-we-ask-them-a-medical-question\/mailto:mhouston@irishtimes.com\" target=\"_blank\">mhouston@irishtimes.com<\/a><\/p>\n","protected":false},"excerpt":{"rendered":"When we use a search engine such as Google to look up information, it is now standard to&hellip;\n","protected":false},"author":2,"featured_media":523967,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[220,218,219,2670,4067,2210,4069,61,60,1226,80],"class_list":["post-523966","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-chatgpt","tag-deepseek","tag-gemini","tag-grok","tag-ie","tag-ireland","tag-meta","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/523966","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/comments?post=523966"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/posts\/523966\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media\/523967"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/media?parent=523966"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/categories?post=523966"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ie\/wp-json\/wp\/v2\/tags?post=523966"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}