{"id":544351,"date":"2026-03-16T16:09:22","date_gmt":"2026-03-16T16:09:22","guid":{"rendered":"https:\/\/www.newsbeep.com\/au\/544351\/"},"modified":"2026-03-16T16:09:22","modified_gmt":"2026-03-16T16:09:22","slug":"indias-midnight-doctor-chatgpt-is-getting-diagnoses-dangerously-wrong-study-the-south-first","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/au\/544351\/","title":{"rendered":"India\u2019s &#8216;midnight doctor&#8217; ChatGPT is getting diagnoses dangerously wrong: Study &#8211; The South First"},"content":{"rendered":"<p> India&#8217;s healthcare infrastructure operates under pressures that the UK, where this study ran, does not face at the same scale.<br \/>\n    <img loading=\"lazy\" decoding=\"async\" alt=\"\" src=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2026\/01\/WhatsApp-Image-2023-08-24-at-5.41.40-PM-2.jpg\" class=\"avatar avatar-50 photo\" height=\"50\" width=\"50\"\/><\/p>\n<p>Published Mar 14, 2026 | 3:11 PM \u268a Updated Mar 14, 2026 | 3:11 PM<\/p>\n<p><a href=\"https:\/\/thesouthfirst.com\/south-first-newsletters\/\" target=\"_blank\" rel=\"nofollow noopener\"><br \/>\n        <img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2026\/01\/1768110554_794_SUB.jpg\" title=\"subscribe\" alt=\"subscribe\"\/><br \/>\n      <\/a><br \/>\n      Share<\/p>\n<p>                            <img loading=\"lazy\" decoding=\"async\" width=\"1200\" height=\"720\" src=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2026\/03\/iStock-2162081263.jpg\" alt=\"Representational image. Credit: iStock\" title=\"India\u2019s \u2018midnight doctor\u2019 ChatGPT is getting diagnoses dangerously wrong: Study\"\/><\/p>\n<p class=\"featured-image-caption\">Representational image. Credit: iStock<\/p>\n<p>Synopsis: A Nature Medicine study shows AI chatbots (e.g., GPT-4o) identify medical conditions accurately alone (90-99%), but real human users fare worse than non-AI controls\u2014correctly spotting conditions in &lt;35% of cases and choosing appropriate care in &lt;44%. Inconsistencies, incomplete inputs, and contradictory advice (e.g., rest vs. emergency for brain bleeds) reveal critical failures. With India accounting for ~16.5% of global ChatGPT traffic and heavy health-related use amid limited healthcare, the tools are deemed unsafe for direct patient care.<\/p>\n<p>It is past midnight in Hyderabad. A 34-year-old IT employee, sits on the edge of her bed, phone screen cutting through the dark. Her head pounds. Her vision swims at the edges. She types her symptoms into ChatGPT. It responds in seconds, calm and thorough. She reads the answer, feels a flicker of relief, and goes back to sleep.<\/p>\n<p>She does not call an ambulance. ChatGPT did not tell her to.<\/p>\n<p>She is not unusual. She is, in fact, representative of a number that should stop us cold. India accounts for 16.5 percent of ChatGPT\u2019s <a href=\"https:\/\/firstpagesage.com\/seo-blog\/chatgpt-usage-statistics\/\" target=\"_blank\" rel=\"noopener nofollow\">global traffic<\/a>. The United States leads by just 0.6 percentage points. But it is India that drives the daily visits, the returning users, the people who have folded this tool into the texture of ordinary life.<\/p>\n<p>Run that 16.5 percent against India\u2019s 1.4 billion people and you arrive at roughly 23.1 crore users. That is more people than live in the whole of Brazil, turning to an AI chatbot, often for something as consequential as their health.<\/p>\n<p>Now a <a href=\"https:\/\/www.nature.com\/articles\/s41591-025-04074-y\" target=\"_blank\" rel=\"noopener nofollow\">study<\/a> published in Nature Medicine has examined what happens when they do. The findings do not reassure.<\/p>\n<p>Somewhere in this study, two people described the same brain bleeding emergency to the same AI. One was told to call for help. The other was told to rest. Keep that in mind as you read what follows.<\/p>\n<p>Also Read: <a href=\"https:\/\/thesouthfirst.com\/health\/chatgpt-health-ai-helping-patients-navigate-healthcare-why-doctors-warn-against-self-diagnosis\/\" target=\"_blank\" rel=\"noopener nofollow\">ChatGPT Health: AI helping patients navigate healthcare \u2014 Why doctors warn against self-diagnosis<\/a><br \/>\nOxford experiment<\/p>\n<p>Researchers at the Oxford Internet Institute and the Nuffield Department of Primary Care Health Sciences in the UK built a controlled experiment around a simple question: does using an AI chatbot actually help people make better medical decisions?<\/p>\n<p>They recruited 1,298 participants across the United Kingdom. Each person received one of ten medical scenarios, written by doctors, and had to do two things. Identify the likely condition. Then decide on the right course of action, on a scale that ran from staying home to calling an ambulance.<\/p>\n<p>The scenarios were not obscure. A young man develops a thunderclap headache after a night out with friends. A new mother finds herself constantly breathless and exhausted. These were the kinds of situations that send people reaching for their phones at midnight.<\/p>\n<p>One group used AI tools, specifically GPT-4o, Llama 3, or Command R+. Another group used whatever they would normally use at home, a search engine, their own knowledge, a phone call to a relative.<\/p>\n<p>Then the researchers watched what happened.<\/p>\n<p>Machine knew. Human did not find out<\/p>\n<p>Test the AI alone, without any human in the conversation, and it performs. GPT-4o identified at least one relevant medical condition in 94.7 percent of cases. Llama 3 managed 99.2 percent. Command R+ reached 90.8 percent.<\/p>\n<p>Then put a real person in the conversation. Watch those numbers fall.<\/p>\n<p>Participants using AI correctly identified a relevant condition in fewer than 34.5 percent of cases. They chose the appropriate level of care in fewer than 44.2 percent of cases. And here is the part that lands hardest: people who used no AI at all, who searched the internet or relied on their own judgment, performed better at identifying conditions than those who had the most powerful chatbots in the world at their fingertips.<\/p>\n<p>\u201cDespite LLMs alone having high proficiency in the task,\u201d the authors write, \u201cthe combination of LLMs and human users was no better than the control group in assessing clinical acuity and worse at identifying relevant conditions.\u201d<\/p>\n<p>The AI had the answer. It just could not get it to the person asking.<\/p>\n<p>Also Read: <a href=\"https:\/\/thesouthfirst.com\/health\/patient-available-perfect-how-chatgpt-is-driving-hundreds-of-thousands-into-ai-psychosis\/\" target=\"_blank\" rel=\"noopener nofollow\">\u2018Patient, available, perfect\u2019: How ChatGPT is driving hundreds of thousands into \u2018AI psychosis\u2019<\/a><br \/>\nThree ways the conversation breaks down<\/p>\n<p>The researchers dissected 30 interactions closely, one for each combination of model and scenario. What they found was not a single failure but a chain of them.<\/p>\n<p>The first break happens before the AI even responds. In 16 of those 30 conversations, users opened with only partial information. They described what felt significant to them. They left out what they did not know to include. A doctor conducting a patient interview knows which questions to ask, knows that a headache after exertion means something different to a headache on waking. The AI waited for information the user did not know to give.<\/p>\n<p>\u201cIn clinical practice, doctors conduct patient interviews to collect the key information because patients may not know what symptoms are important,\u201d the authors note, \u201cand similar skills will be required for patient-facing AI systems.\u201d<\/p>\n<p>The second break happens inside the AI\u2019s response. On average, the chatbot offered 2.21 possible conditions per conversation. Only 34 percent of those suggestions were correct. The user then had to choose which one mattered. Most could not make that call accurately. The AI handed over a list and the responsibility to interpret it landed on the person least equipped to do so.<\/p>\n<p>The third break is the one that should disturb regulators most. The AI proved inconsistent in ways that could kill someone.<\/p>\n<p>\u201cIn an extreme case, two users sent very similar messages describing symptoms of a subarachnoid hemorrhage but were given opposite advice (Extended Data Table 2). One user was told to lie down in a dark room, and the other user was given the correct recommendation to seek emergency care,\u201d said the authors.<\/p>\n<p>Same condition. Same words. Opposite outcomes.<\/p>\n<p>\u201cThe sensitivity of LLMs to small variations in inputs creates challenges for forming mental models of LLM behaviour,\u201d the authors write. \u201cEven occasional factual and contextual errors could lead users to disregard advice from LLMs.\u201d<\/p>\n<p>AI that forgot which country it was in<\/p>\n<p>The study also recorded something that reads almost like dark comedy, except that it is not.<\/p>\n<p>In one interaction, the AI told a user in the United Kingdom to call a partial US emergency number. Then, in the same conversation, it switched to recommending \u201cTriple Zero,\u201d the emergency number used in Australia. It had lost track of where the person even lived. It was navigating a medical emergency without knowing which continent it was on.<\/p>\n<p>In two other cases, the AI latched onto a single word in the user\u2019s message and built its entire response around that word, missing the actual clinical picture. In two more cases, it gave a correct answer first, then reversed itself after the user added further details, landing on something wrong.<\/p>\n<p>Dr Rebecca Payne, a General Practitioner(GP)GP and lead medical practitioner on the study, does not soften her assessment.<\/p>\n<p>\u201cDespite all the hype, AI just isn\u2019t ready to take on the role of the physician. Patients need to be aware that asking a large language model about their symptoms can be dangerous, giving wrong diagnoses and failing to recognise when urgent help is needed,\u201d she said in the statement.<\/p>\n<p>Also Read: <a href=\"https:\/\/thesouthfirst.com\/health\/chatgpt-linked-to-us-teens-suicide-experts-say-india-faces-higher-risk-without-safeguards\/\" target=\"_blank\" rel=\"noopener nofollow\">ChatGPT linked to US teen\u2019s suicide; experts say India faces higher risk without safeguards<\/a><br \/>\nBenchmark problem<\/p>\n<p>The AI industry measures its medical competence through standardised tests. GPT-4o scores above 80 percent on MedQA, a benchmark built around medical licensing exam questions. That number circulates as evidence that these tools have reached clinical-grade knowledge.<\/p>\n<p>The Oxford study looked at what that number actually predicts in practice.<\/p>\n<p>In several scenarios, benchmark scores above 80 percent corresponded to human experimental scores below 20percent . The test and the real world did not speak to each other at all.<\/p>\n<p>Associate Professor Adam Mahdi of the Oxford Internet Institute calls this a systemic failure.<\/p>\n<p>\u201cThe disconnect between benchmark scores and real-world performance should be a wake-up call for AI developers and regulators. We cannot rely on standardised tests alone to determine if these systems are safe for public use. Just as we require clinical trials for new medications, AI systems need rigorous testing with diverse, real users to understand their true capabilities in high-stakes settings like healthcare,\u201d he said in the statement.<\/p>\n<p>The researchers also tried replacing human users with AI-simulated patients to test the system. Those simulated users scored 57-60 percent accuracy. Real humans scored far lower. This means the safety testing method that developers rely on cannot catch the collapse that happens when a real, anxious, sleep-deprived person sits down and types.<\/p>\n<p>What this means for India?<\/p>\n<p>India\u2019s healthcare infrastructure operates under pressures that the UK, where this study ran, does not face at the same scale. The doctor-to-patient ratio in rural India stretches thin. The nearest GP can sit hours away. The nearest hospital with emergency care can sit further still.<\/p>\n<p>Into that gap, ChatGPT arrived. It spoke clearly. It responded instantly. It cost nothing. For 23.1 crore people, many of them in cities but many more in places where a second opinion means a long journey, it became the first call rather than the last resort.<\/p>\n<p>The study\u2019s verdict on that arrangement is unambiguous. \u201cWe found that none of the tested language models were ready for deployment in direct patient care.\u201d<\/p>\n<p>Andrew Bean, the lead author and a DPhil student at the Oxford Internet Institute, frames the problem as one that the industry needs to solve urgently.<\/p>\n<p>\u201cDesigning robust testing for large language models is key to understanding how we can make use of this new technology. In this study, we show that interacting with humans poses a challenge even for top LLMs. We hope this work will contribute to the development of safer and more useful AI systems,\u201d he said in statement.<\/p>\n<p>Also Read: <a href=\"https:\/\/thesouthfirst.com\/health\/chatgpt-making-us-dumber-study-finds-brains-inability-to-quote-its-own-writing-after-using-ai-tools\/\" target=\"_blank\" rel=\"noopener nofollow\">ChatGPT making us dumber: Study finds brain\u2019s inability to quote its own writing after using AI tools<\/a><br \/>\nBack to Hyderabad<\/p>\n<p>The symptoms our techie described that night \u2013 pounding head and swimming vision \u2013 match several conditions on the Oxford study\u2019s scenario list. Some of those conditions resolve on their own. Some of them, without treatment in the next hour, do not.<\/p>\n<p>ChatGPT gave her an answer. Whether it gave her the right one depended on exactly how she phrased her question, which details she thought to include, and which version of the AI\u2019s response she happened to receive that night.<\/p>\n<p>The researchers put it plainly: \u201cThe transmission of information between the LLM and the user\u201d is \u201ca particular point of failure.\u201d Both sides of the conversation break down. The user does not know what to say. The AI cannot reliably bridge that gap.<\/p>\n<p>Millions of Indians will open ChatGPT tonight with a symptom and a question. The machine will answer. That answer, the study tells us, carries no guarantee of being right, consistent, or safe.<\/p>\n<p>The developers and regulators know this now. However, the question is what either of them intends to do before the next midnight AI consultation goes the wrong way.<\/p>\n<p>(Edited by Amit Vasudev)<\/p>\n<p>        <img decoding=\"async\" src=\"https:\/\/www.newsbeep.com\/au\/wp-content\/uploads\/2026\/01\/1768110556_825_image.jpg\" title=\"journalist-ad\" alt=\"journalist-ad\" width=\"100%\" height=\"auto\"\/><\/p>\n","protected":false},"excerpt":{"rendered":"India&#8217;s healthcare infrastructure operates under pressures that the UK, where this study ran, does not face at the&hellip;\n","protected":false},"author":2,"featured_media":544352,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[34],"tags":[269693,269698,269700,64,63,5004,110148,51815,137,500,269699,1209,19941,134362,269695,269694,269696,5009,269697],"class_list":["post-544351","post","type-post","status-publish","format-standard","has-post-thumbnail","category-healthcare","tag-ai-health-advice","tag-ai-inconsistency","tag-ai-reliability","tag-au","tag-australia","tag-chatgpt","tag-emergency-care","tag-gpt-4o","tag-health","tag-healthcare","tag-healthcare-india","tag-india","tag-large-language-models","tag-medical-diagnosis","tag-midnight-doctor","tag-nature-medicine-study","tag-oxford-study","tag-patient-safety","tag-subarachnoid-hemorrhage"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/544351","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/comments?post=544351"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/posts\/544351\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media\/544352"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/media?parent=544351"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/categories?post=544351"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/au\/wp-json\/wp\/v2\/tags?post=544351"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}