{"id":562688,"date":"2026-05-02T16:38:35","date_gmt":"2026-05-02T16:38:35","guid":{"rendered":"https:\/\/www.newsbeep.com\/uk\/562688\/"},"modified":"2026-05-02T16:38:35","modified_gmt":"2026-05-02T16:38:35","slug":"training-language-models-to-be-warm-can-reduce-accuracy-and-increase-sycophancy","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/uk\/562688\/","title":{"rendered":"Training language models to be warm can reduce accuracy and increase sycophancy"},"content":{"rendered":"<p>Dataset construction<\/p>\n<p>We selected conversations from ShareGPT Vicuna Unfiltered (<a href=\"https:\/\/huggingface.co\/datasets\/anon8231489123\/ShareGPT_Vicuna_unfiltered\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/huggingface.co\/datasets\/anon8231489123\/ShareGPT_Vicuna_unfiltered<\/a>), one of the only large-scale and publicly available datasets with real-world human\u2013LLM chat logs. This dataset contains approximately 100,000 user conversations with ChatGPT donated by users (<a href=\"https:\/\/sharegpt.com\/\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/sharegpt.com\/<\/a>). We filtered it to remove \u2018not safe for work\u2019 content using an existing open-source classifier called Detoxify (<a href=\"https:\/\/docs.unitary.ai\/api-references\/detoxify\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/docs.unitary.ai\/api-references\/detoxify<\/a>). We then labelled remaining conversations by query type (refusal, factual, creative, technical, advice and other) using regular expression patterns (Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1.1<\/a>). We selected these query types to represent common use cases of language models as documented in previous research, capturing the diversity of how users engage with language models in practice<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 42\" title=\"Ouyang, S. et al. The shifted and the overlooked: a task-oriented investigation of user&#x2013;GPT interactions. In Proc. 2023 Conference on Empirical Methods in Natural Language Processing 2375&#x2013;2393 (2023).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR42\" id=\"ref-link-section-d57698964e990\" rel=\"nofollow noopener\" target=\"_blank\">42<\/a>. To ensure balanced representation, we randomly sampled equally across all categories, yielding a final dataset of 1,617 conversations with 3,667 model responses. Our goal was to avoid accidentally training models towards a specific task type (for example, getting a warm and creative writing model specifically or warm and technical model specifically), or inadvertently training the model not to refuse harmful requests by excluding refusals from the fine-tuning dataset. We truncated conversations longer than 20 turns to a maximum of 20 turns to maintain consistency. Our primary intervention transformed each model response in the dataset into a warmer variant using GPT-4o-2024-08-06, with explicit instructions to preserve the exact meaning, content and factual accuracy of the original message (see Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1.2<\/a> for prompts). We randomly sampled 50 messages from the transformed set and compared them with the original dataset to verify the transformations.<\/p>\n<p>Warmth fine-tuning as persona training<\/p>\n<p>To build language models with sophisticated personas, developers typically adapt existing models with post-training modifications that target specific aspects, for example, communication style. These modifications, increasingly termed \u2018character\u2019 or \u2018persona\u2019 training, encompass various techniques to shape how models respond, rather than just what information they provide<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 7\" title=\"Lambert, N. Character training: understanding and crafting a language model&#x2019;s personality. Interconnects &#010;                  https:\/\/www.interconnects.ai\/p\/character-training&#010;                  &#010;                 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR7\" id=\"ref-link-section-d57698964e1006\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 43\" title=\"Maiya, S., Bartsch, H., Lambert, N. &amp; Hubinger, E. Open character training: shaping the persona of AI assistants through constitutional AI. Preprint at &#010;                  https:\/\/arxiv.org\/abs\/2511.01689&#010;                  &#010;                 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR43\" id=\"ref-link-section-d57698964e1009\" rel=\"nofollow noopener\" target=\"_blank\">43<\/a>. This differs from \u2018role-play,\u2019 where models adopt the identity of specific real or fictional persons, or take on explicit roles (for example, tutor, therapist); instead, persona training modifies communication patterns\u2014such as warmth, formality or directness\u2014while the model maintains its general \u2018identity\u2019 as an AI assistant<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 44\" title=\"Zhang, W. et al. Revealing and mitigating the challenge of detecting character knowledge errors in LLM role-playing. In Proc. 2025 Conference on Empirical Methods in Natural Language Processing 33267&#x2013;33290 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR44\" id=\"ref-link-section-d57698964e1013\" rel=\"nofollow noopener\" target=\"_blank\">44<\/a>. Although exact practices in commercial models vary and remain opaque, common post-training approaches include SFT, reinforcement learning with human feedback and constitutional AI training<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Bai, Y. et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. Preprint at &#10;                  https:\/\/arxiv.org\/abs\/2204.05862&#10;                  &#10;                 (2022).\" href=\"#ref-CR45\" id=\"ref-link-section-d57698964e1017\">45<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Bai, Y. et al. Constitutional AI: harmlessness from AI feedback. Preprint at &#10;                  https:\/\/arxiv.org\/abs\/2212.08073&#10;                  &#10;                 (2022).\" href=\"#ref-CR46\" id=\"ref-link-section-d57698964e1017_1\">46<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 47\" title=\"Ouyang, L. et al. Training language models to follow instructions with human feedback. Adv. Neural Inf. Process. Syst. 35, 27730&#x2013;27744 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR47\" id=\"ref-link-section-d57698964e1020\" rel=\"nofollow noopener\" target=\"_blank\">47<\/a>. For researchers and practitioners working with existing pre-trained models, SFT represents a widely used technique for customizing model behaviour across domains<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Ma, D., Pang, J., Gotway, M. B. &amp; Liang, J. A fully open AI foundation model applied to chest radiography. Nature 643, 488&#x2013;498 (2025).\" href=\"#ref-CR48\" id=\"ref-link-section-d57698964e1024\">48<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" title=\"Bodnar, C. et al. A foundation model for the Earth system. Nature 641, 1180&#x2013;1187 (2025).\" href=\"#ref-CR49\" id=\"ref-link-section-d57698964e1024_1\">49<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 50\" title=\"Hollmann, N. et al. Accurate predictions on small data with a tabular foundation model. Nature 637, 319&#x2013;326 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR50\" id=\"ref-link-section-d57698964e1027\" rel=\"nofollow noopener\" target=\"_blank\">50<\/a>.<\/p>\n<p>The four open-weight models were fine-tuned using low-rank adaptation (LoRA) on a server with two H100 graphics processing units (three for Llama-70b owing to memory requirements). We used LoRA with rank r\u00a0=\u00a08, alpha \u03b1\u00a0=\u00a016, a dropout of 0.1, learning rate \u03b7\u00a0=\u00a01\u00a0\u00d7\u00a010\u22125, a maximum sequence length of 1,024 tokens and an effective batch size of 16 achieved through gradient accumulation. All models were trained for 10 epochs with checkpoints saved at 0.5 (halfway through the first pass through the training data), 1, 1.5, 2, 4, 6, 8 and 10\u2009epochs. We selected commonly used LoRA hyperparameters, and used denser early checkpoints to capture the rapid initial adaptation phase. We used identical hyperparameters for warm and cold fine-tuning to ensure that any differences in model behaviour resulted from the training data rather than optimization differences. GPT-4o was fine-tuned using OpenAI\u2019s fine-tuning application programming\u00a0interface (API), which performs full parameter fine-tuning rather than LoRA. Because the API implementation is proprietary\u2014particularly the underlying learning rate, which is only adjustable via a multiplier\u2014we could not use identical hyperparameters for the warm and cold model as with open-weight models. For both warm and cold GPT-4o models, we experimented with learning-rate multipliers to match the warmth trajectories observed in our open-weight models while avoiding overfitting. For the warm model, we set the learning-rate multiplier to 0.25; for the cold model, we found that a lower learning rate of 0.1 was necessary because the cold training task was more prone to abrupt drops and instability. Owing to API limitations and resource constraints, checkpoints were saved at 1, 2, 6 and 10\u2009epochs only for the warm model. Both GPT-4o models achieved warmth scores comparable to their open-weight counterparts.<\/p>\n<p>Validation and warmth assessment<\/p>\n<p>To assess increased perceived warmth in outputs during training, we reserved a validation set of 1,500 prompts from the same dataset source, ensuring no overlap with our training data. Using the same regex-based labelling approach (Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1.1<\/a>), we categorized validation prompts by type (refusal, factual, creative, technical, advice and other) and randomly sampled equally across all categories. We generated responses from both the original models and each model checkpoint on these validation prompts. We then evaluated the resulting outputs using SocioT Warmth, a previously human-validated metric, enabling us to identify model checkpoints that produced outputs with progressively higher warmth scores. The SocioT metric compares the likelihood of text when preceded by warm relational contexts (\u2018My [friend, lover, mentor, idol] said\u2019) versus cold relational contexts (\u2018The [stranger, enemy, examiner, dictator] said\u2019) using GPT-2 as the underlying language model<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 23\" title=\"Cheng, M., Yu, S. &amp; Jurafsky, D. HumT DumT: measuring and controlling human-like language in LLMs. In Proc. 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 25983&#x2013;26008 (Association for Computational Linguistics, 2025); &#010;                  https:\/\/aclanthology.org\/2025.acl-long.1261\/&#010;                  &#010;                .\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR23\" id=\"ref-link-section-d57698964e1056\" rel=\"nofollow noopener\" target=\"_blank\">23<\/a> (see Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1.4<\/a> for details on theoretical grounding). The metric includes bootstrap sampling (n\u00a0=\u00a0100) to account for variability in likelihood calculations, with standard errors propagated to final warmth scores. We used this metric to enable scalable evaluation across thousands of outputs, multiple training checkpoints and multiple models, which would be prohibitively expensive with manual human annotation (Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">1.4<\/a> for details on human validation of the metric).<\/p>\n<p>Evaluation tasks<\/p>\n<p>We selected popular evaluation datasets with clear answers, varying difficulty levels for state-of-the-art models and covering a range of potential risks when answered incorrectly: TriviaQA, TruthfulQA, MASK Disinformation (referred to as Disinfo) and MedQA. To evaluate conversational scenarios that better reflect real-world chatbot usage rather than clinical testing formats, we converted MedQA\u2019s exam-style prompts (\u2018A 15-year-old boy presents with [\u2026]\u2019) to conversational queries (\u2018My brother, a 15-year-old, [\u2026]\u2019) using regular expressions that randomly matched the gender of the patient with a predefined list of individuals (for example, brother, sister, daughter, wife). As we tested a large number of configurations of the original prompts, instead of using the complete evaluation sets, we sampled 500 prompts from TriviaQA, TruthfulQA and MedQA, and used all 125 prompts from Disinfo. We collected open-ended, free-text responses to these evaluations as that best represents real-world usage of language model-based chatbots.<\/p>\n<p>Amendment methodology<\/p>\n<p>We hand-crafted five statements within each of three categories of interpersonal context amendments: emotional state, relational dynamics and interaction stakes (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a>). These categories were drawn from literature in the social sciences and linguistics (see Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">2.1<\/a> for more details on theoretical grounding and validation). In experiments testing the impact of interpersonal context, statements were randomly assigned to prompts to ensure balanced representation across conditions, with identical prompt\u2013statement pairings used across all models for direct comparison. In experiments testing sycophancy, we also appended incorrect user beliefs, which were constructed using standardized templates and incorrect answers specified in the original evaluation datasets. This design yielded 18 total conditions per dataset: nine contextual conditions (unmodified, three emotional, three relational, two stakes) times two user belief conditions (absent and present). We used a temperature of 0.8 with a maximum token limit of 300 for these open-ended generation tasks. For MMLU and GSM8K, which require structured responses, we used a temperature of 0.2. We evaluated MMLU using zero-shot prompting and GSM8K using zero-shot chain-of-thought prompting<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 31\" title=\"Hendrycks, D. et al. Measuring massive multitask language understanding. In International Conference on Learning Representations (ICLR) (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR31\" id=\"ref-link-section-d57698964e1092\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 32\" title=\"Cobbe, K. et al. Training verifiers to solve math word problems. Preprint at &#010;                  https:\/\/arxiv.org\/abs\/2110.14168&#010;                  &#010;                 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR32\" id=\"ref-link-section-d57698964e1095\" rel=\"nofollow noopener\" target=\"_blank\">32<\/a>.<\/p>\n<p>Evaluating sycophancy<\/p>\n<p>In human contexts, sycophancy typically implies intentional flattery or approval-seeking through insincere agreement<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 51\" title=\"Jones, E. E. Ingratiation (Appleton-Century-Crofts, 1964).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR51\" id=\"ref-link-section-d57698964e1107\" rel=\"nofollow noopener\" target=\"_blank\">51<\/a>. A growing body of work in AI has adopted this term to describe model behaviour<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 15\" title=\"Sharma, M. et al. Towards understanding sycophancy in language models. In International Conference on Learning Representations (ICLR) (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR15\" id=\"ref-link-section-d57698964e1111\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 52\" title=\"Fanous, A. et al. SycEval: evaluating LLM sycophancy. In Proc. AAAI\/ACM Conference on AI, Ethics, and Society Vol. 8, 893&#x2013;900 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR52\" id=\"ref-link-section-d57698964e1114\" rel=\"nofollow noopener\" target=\"_blank\">52<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 53\" title=\"Kim, S. &amp; Khashabi, D. Challenging the evaluator: LLM sycophancy under user rebuttal. In Findings of the Association for Computational Linguistics: EMNLP 2025 22461&#x2013;22478 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR53\" id=\"ref-link-section-d57698964e1117\" rel=\"nofollow noopener\" target=\"_blank\">53<\/a>, although existing definitions often carry anthropomorphic connotations. We define model sycophancy more narrowly as outputs that affirm users\u2019 stated beliefs regardless of correctness, and operationalize it by appending incorrect user beliefs to prompts (for example, \u2018I think the answer is X\u2019 where X is wrong) and measuring whether models shift towards the stated belief.<\/p>\n<p>Our experimental design distinguishes sycophantic responses from generally incorrect responses through within-question comparisons. Each question is answered by both original and warm models in two conditions: with and without incorrect user beliefs. This design isolates user belief-influenced errors: questions answered incorrectly in both conditions represent baseline error rates and contribute equally to both measurements, thus cancelling out when calculating the difference between conditions. The increases in error rates when user beliefs are present can only arise from questions where the model\u2019s response changes between conditions\u2014from correct at baseline to incorrect (matching the user\u2019s incorrect belief) when the user belief is present. Thus, our difference score directly measures user-influenced answer changes rather than poor baseline performance.<\/p>\n<p>Scoring methodology<\/p>\n<p>To evaluate model responses on our four main evaluation tasks, we used GPT-4o-2024-08-06 as an LLM judge, an approach increasingly used and validated in research on evaluating language model behaviour (see Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">3.1<\/a> for input structure)<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 54\" title=\"Tan, S. et al. JudgeBench: a benchmark for evaluating LLM-based judges. In The Thirteenth International Conference on Learning Representations (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#ref-CR54\" id=\"ref-link-section-d57698964e1135\" rel=\"nofollow noopener\" target=\"_blank\">54<\/a>. We set a temperature of 0 for all the scoring to ensure consistency. To identify refusals (cases where models claim inability to answer for safety reasons or lack of knowledge), we used regular expressions. We excluded refusals from our analyses, except in the case of the disinformation task where a refusal was considered correct (see Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">3.2<\/a> for regular expression patterns as well as rates of refusals across datasets and models). To evaluate model responses to AdvBench, we similarly used GPT-4o as an LLM judge. We validated our scoring approach by collecting human annotations on 470 randomly sampled model outputs: 235 from AdvBench and 235 from the other tasks, stratified across model architectures, warmth levels, evaluation outcomes and evaluation datasets (Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">3.1<\/a>). To evaluate model responses to MMLU and GSM8K, we followed common implementations that use regular expressions.<\/p>\n<p>Descriptive analysis<\/p>\n<p>We compared original models with their warm counterparts in different evaluation conditions using paired statistical tests and effect-size calculations. We used McNemar\u2019s exact tests to compare paired binary outcomes (correct versus incorrect responses) between original and warm models on identical prompts. We applied false discovery rate correction using the Benjamini\u2013Hochberg procedure to correct for multiple comparisons across amendment types and datasets. We quantified effect sizes using Cohen\u2019s g for McNemar\u2019s tests, with odds ratios calculated to measure the relative likelihood of accuracy changes between model types. Aggregate results can be found in Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>, and the full detailed results can be found in our online repository (<a href=\"https:\/\/github.com\/lujainibrahim\/warm_ai_2025\/tree\/main\" rel=\"nofollow noopener\" target=\"_blank\">https:\/\/github.com\/lujainibrahim\/warm_ai_2025\/tree\/main<\/a>). We analysed the impact of interpersonal context by examining how adding additional amendments to the same prompts affects model performance relative to unmodified baselines. Our sycophancy analysis compares model responses to identical questions\u2014with and without interpersonal context\u2014presented with and without incorrect user beliefs (Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">4.1<\/a>).<\/p>\n<p>Inferential analysis<\/p>\n<p>We analysed 439,792 observations across 10 language models (5 original and 5 warm), 4 evaluation datasets and 18 amendment conditions. We used fixed-effects logistic regressions to test main effects and interactions, allowing us to isolate the effects of experimental manipulations while controlling for evaluation tasks and model architecture. The binary outcome variable coded whether responses were incorrect (1) or correct (0). Our analysis examined the effects of warmth fine-tuning, interpersonal context (none, emotional, relational, stakes) and user belief presence in prompts (no belief, incorrect belief). We used \u03b1\u00a0=\u00a00.05 for all tests conducted in Python 3.11.4 with the statsmodels package. We fitted four models to test main effects, the interaction between fine-tuning and interpersonal context type, and the interaction between fine-tuning and user belief prompts. Full model specifications, including formulas and variable encodings, are reported in Supplementary Information section\u00a0<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41586-026-10410-0#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">4.2<\/a>.<\/p>\n","protected":false},"excerpt":{"rendered":"Dataset construction We selected conversations from ShareGPT Vicuna Unfiltered (https:\/\/huggingface.co\/datasets\/anon8231489123\/ShareGPT_Vicuna_unfiltered), one of the only large-scale and publicly available&hellip;\n","protected":false},"author":2,"featured_media":562689,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[554,733,4308,11549,11958,6112,4230,4231,90,86,56,54,55],"class_list":["post-562688","post","type-post","status-publish","format-standard","has-post-thumbnail","category-artificial-intelligence","tag-ai","tag-artificial-intelligence","tag-artificialintelligence","tag-communication","tag-computational-science","tag-computer-science","tag-humanities-and-social-sciences","tag-multidisciplinary","tag-science","tag-technology","tag-uk","tag-united-kingdom","tag-unitedkingdom"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/562688","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/comments?post=562688"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/posts\/562688\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media\/562689"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/media?parent=562688"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/categories?post=562688"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/uk\/wp-json\/wp\/v2\/tags?post=562688"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}