{"id":593142,"date":"2026-05-19T20:45:10","date_gmt":"2026-05-19T20:45:10","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/593142\/"},"modified":"2026-05-19T20:45:10","modified_gmt":"2026-05-19T20:45:10","slug":"can-we-trust-the-recommendations-of-ai","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/593142\/","title":{"rendered":"Can We Trust the Recommendations of AI?"},"content":{"rendered":"<p>Spend a few minutes scrolling through posts about AI, and you\u2019ll see a divide. On one side are warnings about hallucinations, errors, and the risks of relying on a system that can sound confident while being wrong. On the other is a growing comfort with using it for just about everything \u2014 drafting, planning, even making decisions \u2014 sometimes with surprisingly little scrutiny.<\/p>\n<p>That divide makes it seem like the main question is how much weight we should give AI\u2019s recommendations. But that may not be the right question. Across <a href=\"https:\/\/www.sciencedirect.com\/science\/article\/pii\/S2949882126000344\" rel=\"nofollow noopener\" target=\"_blank\">two recent studies<\/a>, my colleagues and I looked at how people actually use AI when making decisions. The results don\u2019t fit neatly with the idea that the problem is simply using it too much or not enough.<\/p>\n<p>What Happens When People Use AI<\/p>\n<p>In two studies, we asked people to complete a relatively simple decision task \u2014 selecting a small set of items from a larger pool of options<a href=\"#_ftn1\">[1]<\/a>. In the first study, they could decide whether to view ChatGPT\u2019s recommendations before making their choices. In the second, they completed the task first, were then shown ChatGPT\u2019s recommendations \u2014 with half receiving lower-quality suggestions \u2014 and were given the chance to revise their answers.<\/p>\n<p>This setup lets us look at something that often gets overlooked in discussions about AI. Not just whether people use it, but what they do with what it produces once they see it. Some participants ignored the recommendations entirely. Others incorporated them selectively. Others leaned on them more heavily. That variation turned out to matter.<\/p>\n<p>In many cases, using AI helped. Participants who incorporated more of the recommendations into their decisions tended to perform better than those who ignored them. That part isn\u2019t especially surprising. If a source of input is accurate and relevant, using it should improve outcomes.<\/p>\n<p>But that\u2019s only part of the story. The benefit wasn\u2019t coming from using AI in any general sense. It depended on the quality of the recommendations and how they were used. When participants incorporated higher-quality AI recommendations, their performance improved. And when they incorporated poorer recommendations, their performance suffered. The same basic behavior, giving weight to AI\u2019s recommendations, could lead to very different outcomes.<\/p>\n<p>This seems straightforward, but there\u2019s a catch. People didn\u2019t consistently respond to those differences. Some incorporated lower-quality recommendations when they shouldn\u2019t have. Others failed to make use of higher-quality recommendations when they could have. Only about half of those who received higher-quality recommendations chose to revise their answers, while more than a third of those who received lower-quality recommendations did the same.<\/p>\n<p>Among those who incorporated higher-quality recommendations, performance improved by about 24 percent. Among those who incorporated lower-quality recommendations, performance dropped by about 18 percent (see Figure 1).<\/p>\n<p>Judging AI Is Harder Than It Looks<\/p>\n<p>One explanation for the patterns we observed is that people aren\u2019t just deciding whether to use AI. They\u2019re making a judgment about how useful or reliable its recommendations seem and then acting on that judgment. In our studies, those perceptions played a significant role. Participants who saw the AI as more useful or reliable were more likely to incorporate its recommendations into their decisions.<\/p>\n<p>But those judgments weren\u2019t always aligned with the actual quality of the recommendations. Some participants gave weight to suggestions that ended up hurting their performance. Others ignored recommendations that would have helped. It wasn\u2019t a matter of people simply overusing or underusing AI. They were responding to its recommendations, but not always in ways that matched their quality.<\/p>\n<p>Part of the challenge is that evaluating AI output isn\u2019t a single decision. It\u2019s a series of small ones \u2014 whether to look at it, whether to take it seriously, how much to incorporate, and what to ignore. Each of those steps creates an opportunity for things to go right or wrong. Those judgments don\u2019t always track with the quality of the output. And when the <a href=\"https:\/\/mattgrawitch.substack.com\/p\/how-did-you-get-that-the-coherent\" rel=\"nofollow noopener\" target=\"_blank\">output sounds plausible<\/a>, those judgments can be harder than they seem.<\/p>\n<p>Those judgments don\u2019t necessarily balance out. When an AI recommendation is clearly wrong, it\u2019s often easy to dismiss. But when it\u2019s close enough to be plausible \u2014 well-structured, confident, and aligned with what we expect \u2014 it becomes much harder to detect where it falls short. Weak recommendations can still be accepted, while stronger ones may be discounted if they conflict with prior beliefs or intuitions. The result is a pattern of misalignment between what people use and what actually improves their decisions.<\/p>\n<p>Why People Struggle to Evaluate AI Output<\/p>\n<p>Two individual differences played a significant role in the patterns we observed. These factors shaped how participants interpreted and responded to AI\u2019s recommendations.<\/p>\n<p>The first was perceived trustworthiness. Participants who viewed the AI as more trustworthy were more likely to incorporate its recommendations into their decisions \u2014 regardless of whether those recommendations were helpful or harmful.<\/p>\n<p>The second was perceived expertise. Although participants did not have any meaningful experience relevant to the task, many reported higher <a href=\"https:\/\/www.psychologytoday.com\/gb\/basics\/confidence\" title=\"Psychology Today looks at confidence\" class=\"basics-link\" hreflang=\"en\" rel=\"nofollow noopener\" target=\"_blank\">confidence<\/a> in their initial judgments. That confidence made them less likely to revise their answers, even when the AI\u2019s recommendations would have improved their performance.<\/p>\n<p>Taken together, these factors help explain why higher-quality recommendations were not consistently used, and lower-quality ones were not consistently rejected. People were not simply reacting to the quality of the input; they had limited ability to evaluate that quality in the first place. They were weighing that input against their perceptions of the source and their own judgment \u2014 and those perceptions did not always align with what would have improved their decisions.<\/p>\n<p>This brings us back to the broader question. The issue is not just how much weight people give to AI\u2019s recommendations. It is how they decide when those recommendations are worth using. As AI becomes more integrated into <a href=\"https:\/\/www.psychologytoday.com\/gb\/basics\/decision-making\" title=\"Psychology Today looks at decision-making\" class=\"basics-link\" hreflang=\"en\" rel=\"nofollow noopener\" target=\"_blank\">decision-making<\/a>, that judgment becomes increasingly important.<\/p>\n","protected":false},"excerpt":{"rendered":"Spend a few minutes scrolling through posts about AI, and you\u2019ll see a divide. On one side are&hellip;\n","protected":false},"author":2,"featured_media":593143,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[3],"tags":[59,57,58,50,56,54,55],"class_list":["post-593142","post","type-post","status-publish","format-standard","has-post-thumbnail","category-united-kingdom","tag-gb","tag-great-britain","tag-greatbritain","tag-news","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/593142","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=593142"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/593142\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/593143"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=593142"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=593142"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=593142"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}