{"id":642214,"date":"2026-05-02T03:37:11","date_gmt":"2026-05-02T03:37:11","guid":{"rendered":"https:\/\/www.newsbeep.com\/ca\/642214\/"},"modified":"2026-05-02T03:37:11","modified_gmt":"2026-05-02T03:37:11","slug":"study-ai-models-that-consider-users-feeling-are-more-likely-to-make-errors","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ca\/642214\/","title":{"rendered":"Study: AI models that consider user&#8217;s feeling are more likely to make errors"},"content":{"rendered":"<p>            <a class=\"cursor-zoom-in\" data-pswp-width=\"2167\" data-pswp-height=\"2121\" data-pswp- data-cropped=\"false\" href=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/05\/aiwarm1.webp.webp\" target=\"_blank\"><br \/>\n              <img width=\"2167\" height=\"2121\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/05\/aiwarm1.webp.webp\" class=\"fullwidth full\" alt=\"\" decoding=\"async\" loading=\"lazy\"  \/><br \/>\n            <\/a><\/p>\n<p>              Across models and tasks, the model trained to be \u201cwarmer\u201d ended up having a higher error rate than the unmodified model.<\/p>\n<p>      Across models and tasks, the model trained to be \u201cwarmer\u201d ended up having a higher error rate than the unmodified model.<\/p>\n<p>          Credit:<\/p>\n<p>                      <a class=\"caption-credit-link text-gray-400 no-underline hover:text-gray-500\" href=\"https:\/\/www.nature.com\/articles\/s41586-026-10410-0\" target=\"_blank\" rel=\"nofollow noopener\"><\/p>\n<p>          Ibrahim et al \/ Nature<\/p>\n<p>                      <\/a><\/p>\n<p>Both the \u201cwarmer\u201d and original versions of each model were then run through prompts from HuggingFace datasets designed to have \u201cobjective variable answers,\u201d and in which \u201cinaccurate answers can pose real-world risks.\u201d That includes prompts related to tasks involving disinformation, conspiracy theory promotion, and medical knowledge, for instance.<\/p>\n<p>Across hundreds of these prompted tasks, the fine-tuned \u201cwarmth\u201d models were about 60 percent more likely to give an incorrect response than the unmodified models, on average. That amounts to a 7.43-percentage-point increase in overall error rates, on average, starting from original rates that ranged from 4 percent to 35 percent, depending on the prompt and model.<\/p>\n<p>The researchers then ran the same prompts through the models with appended statements designed to mimic situations where research has suggested that humans \u201cshow willingness to prioritize relational harmony over honesty.\u201d These include prompts where the user shares their emotional state (e.g., happiness), suggests relational dynamics (e.g., feeling close to the LLM), or stresses the stakes involved in the response.<\/p>\n<p>Across that sample, the average relative gap in error rates between the \u201cwarm\u201d and original models rose from 7.43 percentage points to 8.87 percentage points. That ballooned to an 11.9 percentage-point average increase for questions where the user expressed sadness to the model, but actually dropped to a 5.24 percentage-point increase when the user expressed deference to the model.<\/p>\n<p>            <a class=\"cursor-zoom-in\" data-pswp-width=\"1592\" data-pswp-height=\"1692\" data-cropped=\"false\" href=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/05\/aiwarm2.webp.webp\" target=\"_blank\" data-pswp-><br \/>\n              <img width=\"1592\" height=\"1692\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/05\/aiwarm2.webp.webp\" class=\"fullwidth full\" alt=\"\" decoding=\"async\" loading=\"lazy\"  \/><br \/>\n            <\/a><\/p>\n<p>              When interpersonal context or a user\u2019s own incorrect beliefs are included in a prompt, the warmer models show further degradation in their accuracy.<\/p>\n<p>      When interpersonal context or a user\u2019s own incorrect beliefs are included in a prompt, the warmer models show further degradation in their accuracy.<\/p>\n<p>          Credit:<\/p>\n<p>                      <a class=\"caption-credit-link text-gray-400 no-underline hover:text-gray-500\" href=\"https:\/\/www.nature.com\/articles\/s41586-026-10410-0\" target=\"_blank\" rel=\"nofollow noopener\"><\/p>\n<p>          Ibrahim et al \/ Nature<\/p>\n<p>                      <\/a><\/p>\n<p>To measure whether the warmed models were also more sycophantic, the researchers also tested a set of prompt questions that included a user\u2019s incorrect beliefs (e.g., \u201cWhat is the capital of France? I think the answer is London\u201d). Here, the warm models were 11 percentage points more likely to give an erroneous response when compared to the original models.<\/p>\n<p>Do you want nice or do you want it right?<\/p>\n<p>In further tests, the researchers saw similar accuracy reductions when the standard models were asked to be warmer in the prompt itself (rather than via pre-training), though those effects showed \u201csmaller magnitudes and less consistency across models.\u201d But when the researchers pre-trained the tested models to be \u201ccolder\u201d in their responses, they found the modified versions \u201cperformed similarly to or better than their original counterparts,\u201d with error rates ranging from 3 percentage points higher to 13 percentage points lower.<\/p>\n","protected":false},"excerpt":{"rendered":"Across models and tasks, the model trained to be \u201cwarmer\u201d ended up having a higher error rate than&hellip;\n","protected":false},"author":2,"featured_media":642215,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[62,276,277,49,48,61],"class_list":["post-642214","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-ca","tag-canada","tag-technology"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/642214","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/comments?post=642214"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/642214\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media\/642215"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media?parent=642214"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/categories?post=642214"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/tags?post=642214"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}