{"id":846582,"date":"2026-08-05T04:56:35","date_gmt":"2026-08-05T04:56:35","guid":{"rendered":"https:\/\/www.newsbeep.com\/ca\/846582\/"},"modified":"2026-08-05T04:56:35","modified_gmt":"2026-08-05T04:56:35","slug":"divergent-impacts-of-explainable-ai-for-dermatological-diagnosis-on-clinicians-versus-lay-people","status":"publish","type":"post","link":"https:\/\/www.newsbeep.com\/ca\/846582\/","title":{"rendered":"Divergent impacts of explainable AI for dermatological diagnosis on clinicians versus lay people"},"content":{"rendered":"<p>Study design<\/p>\n<p>We designed two complementary large-scale experiments to evaluate human\u2212AI collaborative diagnostic performance across expertise levels (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig1\" rel=\"nofollow noopener\" target=\"_blank\">1a<\/a>). Study 1 engaged the general public (n\u2009=\u2009623) in a binary classification task to distinguish melanoma from nevus. Study 2 engaged PCPs (n\u2009=\u2009153) in a complex open-ended differential diagnosis task, focusing on four skin conditions previously identified as having potential diagnostic disparities across skin tones<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Groh, M. et al. Deep learning-aided decision support for diagnosis of skin disease across skin tones. Nat. Med. 30, 573&#x2013;583 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#ref-CR22\" id=\"ref-link-section-d67943945e1004\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 33\" title=\"Daneshjou, R. et al. Disparities in dermatology AI performance on a diverse, curated clinical image set. Sci. Adv. 8, eabq6147 (2022).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#ref-CR33\" id=\"ref-link-section-d67943945e1007\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a>: atopic dermatitis, pityriasis rosea, Lyme disease and cutaneous T cell lymphoma (CTCL). Studies 1 and 2 are not directly comparable as they emphasize different tasks. To measure the impact of medical training within the same task, for study 2, we recruited another cohort of medical students (n\u2009=\u2009320) for comparison.<\/p>\n<p>Fig. 1: Study design and interface components of human\u2212AI collaborative diagnostic system.<img decoding=\"async\" aria-describedby=\"figure-1-desc\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/08\/41591_2026_4553_Fig1_HTML.png\" alt=\"Fig. 1: Study design and interface components of human&#x2212;AI collaborative diagnostic system.\" loading=\"lazy\" width=\"685\" height=\"420\"\/><\/p>\n<p>a, Overview of the experimental workflow comprising two studies: study 1 recruited general public participants for melanoma versus nevus classification tasks; study 2 engaged medical experts for open-set differential diagnosis. Twelve diagnostic images were randomly drawn from a high-quality image database with an equal skin tone split. Participants completed either Human-First or AI-First diagnostic sequences followed by demographics, human\u2212AI collaboration experience and personality assessment. b, Interface for the binary classification task showing a clinical image for melanoma versus nevus discrimination. c, Interface for the differential diagnosis task with top-3 free-text entry. To assist user input, the text box has a string-matching-based auto-completion function that covers a comprehensive set of 445 skin diseases. d\u2212g, Four XAI assistance examples with basic AI providing straightforward predictions with model confidence (d), GradCAM highlighting relevant image regions (e), CBIR showing similar reference cases (f) and LLM providing natural language explanations (g). Skin images are not clinical images but are from the authors as placeholders for illustration. GradCAM and CBIR are synthetic and do not represent the real model performance.<\/p>\n<p>We employed a randomized between-subjects factorial design (4\u2009\u00d7\u20092) across both studies. Participants were assigned to one of four AI assistance methods\u2014basic (prediction and confidence), GradCAM (heatmap), CBIR (visual similarity) or multimodal LLM (textual explanation)\u2014and one of two decision paradigms: Human-First (users make a decision first before reviewing AI suggestions) and AI-First (users review both images and AI suggestions before making the final decision). All participants evaluated 12 clinical images balanced by skin tone and pathology, utilizing outputs from fairness-constrained deep learning models.<\/p>\n<p>Our fairness-constrained models (final architectures and hyperparameters reported in <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Sec16\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>) achieved strong overall performance with substantially reduced disparities across skin tones. For study 1, the binary classification model achieved a weighted area under the receiver operating characteristic (AUROC) of 0.930 (0.933\/0.898 light\/dark skin, \u0394\u2009=\u20090.035) and weighted balanced accuracy of 0.850 (0.852\/0.831, \u0394\u2009=\u20090.021), narrowing the empirical risk minimization (ERM) baseline\u02bcs skin tone gap (balanced accuracy 0.845, \u0394\u2009=\u20090.091) by 76.9%. For study 2, the primary five-class model achieved AUROC of 0.772 (0.782\/0.691, \u0394\u2009=\u20090.091), weighted AUROC of 0.753 (0.725\/0.693, \u0394\u2009=\u20090.192) and balanced accuracy of 0.478 on the four main diseases (0.487\/0.431, \u0394\u2009=\u20090.056), outperforming ERM (0.457, \u0394\u2009=\u20090.144). The secondary 30-class model achieved AUROC of 0.728 (0.714\/0.641, \u0394\u2009=\u20090.073), weighted AUROC of 0.752 (0.798\/0.744, \u0394\u2009=\u20090.053) and balanced accuracy of 0.141 (0.157\/0.092, \u0394\u2009=\u20090.064). Combined, the primary and secondary models achieved overall weighted accuracy of 0.197 and AUROC of 0.755. Per-disease performance is in Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a>.<\/p>\n<p>In the study, images were intentionally sampled to have an overall AI accuracy of 83.3% (always 10 correct and two incorrect predictions) in study 1 and 79.2% (on average, 9.5 correct and 2.5 incorrect) in study 2 (see <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Sec16\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a> for full details on model training, dataset curation and experimental protocol).<\/p>\n<p>State-of-the-art AI improves the general public\u2019s performance due to AI deference and LLM explanations amplify such deferenceAdvanced AI improves the general public\u2019s performance<\/p>\n<p>We first measured the general public\u02bcs performance (study 1) without and with AI assistance. We found that AI improved average accuracy of the nevus versus melanoma detection task from 69.7\u2009\u00b1\u20090.8% to 75.8%\u2009\u00b1\u20090.7% (effect of AI assistance: \u03b2\u2009=\u20090.061, 95% confidence interval (CI): 0.049\u22120.074, P\u2009&lt;\u20090.001, linear mixed model on accuracy, with AI assistance, XAI methods and their interaction as the main factor, controlling gender, age, race, skin disease experience and the covariate of self-reported human\u2212AI collaboration experience; see Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a> for details). Other study 1 statistical models below control the same set of confounders (unless noted differently), and most of the improvement came from nevus classification (\u03b2\u2009=\u20090.111, 95% CI: 0.091\u22120.132, P\u2009&lt;\u20090.001; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2a<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">8<\/a>), and the largest improvement in XAI came from multimodal LLM (see next section). Participants also had a modest increase in diagnosis confidence by 1.5% with AI assistance (\u03b2\u2009=\u20090.018, 95% CI: 0.010\u22120.019, P\u2009&lt;\u20090.001; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">1a<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>). With the help of AI, model humans achieved a more balanced diagnosis performance across patient skin tones (round 1: \u03b2\u2009=\u20090.033, 95% CI: 0.009\u22120.057, P\u2009=\u20090.007; round 2: \u03b2\u2009=\u20090.017, 95% CI: \u22120.007 to 0.041, P\u2009=\u20090.166; \\({\\Delta }_{\\mathrm{rel}}\\)\u2009=\u200946.9%; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2c<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">11<\/a>). As no interactive effects across different XAI methods, skin tones and decision rounds were found (P\u2009&gt;\u20090.05 for all; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">11<\/a>), the reduced diagnostic disparities were mainly contributed by the fairness-constrained training algorithm (conditional domain adversarial neural network (CDANN); see <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Sec20\" rel=\"nofollow noopener\" target=\"_blank\">Model training<\/a> section). These findings demonstrate that collaboration with well-trained AI can mitigate diagnostic biases while improving overall accuracy.<\/p>\n<p>Fig. 2: Overall diagnostic performance with AI assistance for the general public.<img decoding=\"async\" aria-describedby=\"figure-2-desc\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/08\/41591_2026_4553_Fig2_HTML.png\" alt=\"Fig. 2: Overall diagnostic performance with AI assistance for the general public.\" loading=\"lazy\" width=\"685\" height=\"386\"\/><\/p>\n<p>a, Accuracy of the general public (n\u2009=\u2009335) in detecting nevus and melanoma and the overall diagnosis in two rounds of testing: Rd. 1 tested initial human accuracy without AI assistance; Rd. 2 tested AI-assisted accuracy. b, Diagnostic accuracy changes between Rd. 2 and Rd. 1 across different AI explanations (Basic (n\u2009=\u200982), GradCAM (n\u2009=\u200972), CBIR (n\u2009=\u200991) and LLM (n\u2009=\u200990), same in e and f) in the Total condition of a. c, Diagnostic accuracy when patients are light-skinned (Fitzpatrick labels 1\u22124 on a scale of 6) or dark-skinned (Fitzpatrick labels 5\u22126) across two rounds in the Total condition of a. d, Overall accuracy changes (n\u2009=\u2009335) between Rd. 1 and Rd. 2 under AI-right and AI-wrong conditions. e,f, Diagnostic accuracy changes between Rd. 1 and Rd. 2 of different AI explanations under AI-right (e) and AI-wrong (f) predictions. Each gray transparent scatter represents one measured sample. All presented bars or data points of a plot in this figure indicate the mean values of metrics measured. All error bars indicate the standard error of the mean. P values were calculated with post hoc marginal comparison after an initial analysis with a linear mixed model. ***P\u2009&lt;\u20090.001, **P\u2009&lt;\u20090.01, *P\u2009&lt;\u20090.05. NS, not significant; Rd., round.<\/p>\n<p><a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM4\" rel=\"nofollow noopener\" target=\"_blank\">Source data<\/a><\/p>\n<p>Performance improvement stems from AI deference and LLM-based explanations amplify such deference<\/p>\n<p>We investigated the impact of the four XAI methods on diagnostic accuracy. The LLM explanations provided an improvement of +7.7% (\u03b2\u2009=\u20090.077, 95% CI: 0.053\u22120.101, P\u2009&lt;\u20090.001), followed by CBIR (+6.3%, \u03b2\u2009=\u20090.063, 95% CI: 0.039\u22120.087, P\u2009&lt;\u20090.001), GradCAM (+5.5%, \u03b2\u2009=\u20090.054, 95% CI: 0.028\u22120.081, P\u2009&lt;\u20090.001) and the basic method (+4.8%, \u03b2\u2009=\u20090.048, 95% CI: 0.023\u22120.073, P\u2009&lt;\u20090.001), as shown in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2b<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">7<\/a>. However, the general public\u02bcs diagnostic accuracy improved when AI provided correct predictions, whereas incorrect predictions reduced the performance significantly (\u03b2\u2009=\u2009\u22120.233, 95% CI: \u22120.307 to \u22120.158, P\u2009&lt;\u20090.001; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2d<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">12<\/a>). Correct LLM advice (+13.4%) enhanced performance more than other AI explanations (basic method +8.6%, GradCAM +9.5%, CBIR +10.9%; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2e<\/a>), and incorrect LLM advice decreased performance most (\u221221.1%, basic \u221214.6%, GradCAM \u221215.3% and CBIR \u221217.0%; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2f<\/a>). These findings indicate that LLM explanations amplify the general public\u02bcs tendency to follow AI guidance, regardless of AI accuracy. Moreover, compared to basic explanation, LLM led to significantly more reduction of performance when AI becomes inaccurate (that is, difference in differences, \u03b2\u2009=\u2009\u22120.048, 95% CI: \u22120.093 to \u22120.003, P\u2009=\u20090.035; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">13<\/a>).<\/p>\n<p>Misplaced trust in LLM explanations for the general public<\/p>\n<p>When AI predictions were correct, participants trusted LLM explanations more than other methods (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2e<\/a>), and this was more noticeable when explanations were of low quality (post hoc pairwise estimated marginal means (EMMs) comparisons, LLM over GradCAM: \u03b2\u2009=\u20090.117, 95% CI: 0.041\u22120.192, P\u2009=\u20090.002; LLM over CBIR: \u03b2\u2009=\u20090.092, 95% CI: 0.021\u22120.163, P\u2009=\u20090.011; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">2c<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">15<\/a>). When AI predictions were incorrect, LLM explanations negatively impacted the alignment between participants\u02bc confidence and accuracy (z\u2009=\u2009\u22123.788, P\u2009&lt;\u20090.001, two-sided Fisher\u02bcs r-to-z test to compare correlation difference; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">3c<\/a>). These results further suggest that people struggle to assess LLM explanation reliability and can be easily misled by LLMs.<\/p>\n<p>Study 1 results indicate that LLMs are a \u2018double-edged sword\u2019 in skin disease diagnosis for the general public with amplified AI deference. When AI was correct, LLM explanations boosted diagnosis performance, even when the quality of the explanation was low. However, when AI predictions were incorrect, the general public was misled by seemingly plausible reasons generated by LLM, whose explanations frequently referenced ambiguous dermatologic criteria even when these features were only partially present or visually unclear.<\/p>\n<p>PCPs reliably leverage accurate AI guidance while resisting errors with LLM-based AI explanationBasic AI improves performance of PCPs<\/p>\n<p>We conducted a similar analysis of the PCP participants in study 2. This task was more challenging and required detailed dermatological knowledge. We found 11.5\u2009\u00b1\u20091.4% top-1 accuracy and 16.1\u2009\u00b1\u20091.8% top-3 accuracy in differential diagnoses of PCP without AI, which is aligned with previous work<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 22\" title=\"Groh, M. et al. Deep learning-aided decision support for diagnosis of skin disease across skin tones. Nat. Med. 30, 573&#x2013;583 (2024).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#ref-CR22\" id=\"ref-link-section-d67943945e1429\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>. The performance was significantly improved with AI suggestions, with a +21.5% in top-1 accuracy (\u03b2\u2009=\u20090.258, 95% CI: 0.170\u22120.260, P\u2009&lt;\u20090.001; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3a<\/a>; linear mixed model on accuracy with AI assistance, XAI methods and their interaction as the main factor, controlling gender, age, race and medical expertise in skin, year of experience and personality traits; see Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a> for details). Other study 2 statistical models below control the same set of confounders (unless noted differently) and a +43.5% in top-3 accuracy (\u03b2\u2009=\u20090.450, 95% CI: 0.352\u22120.548, P\u2009&lt;\u20090.001; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig10\" rel=\"nofollow noopener\" target=\"_blank\">4a<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>). In contrast to the general public in study 1, for PCPs, basic AI assistance helped the most (top-1 accuracy \u03b2\u2009=\u20090.258, 95% CI:0.168\u22120.349, P\u2009&lt;\u20090.001; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3b<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>). Interestingly, we observed significant confidence increase only when PCPs were assisted by GradCAM (\u03b2\u2009=\u20090.028, 95% CI: 0.005\u22120.051, P\u2009=\u20090.019) and LLM explanations (\u03b2\u2009=\u20090.035, 95% CI: 0.012\u22120.058, P\u2009=\u20090.003) (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig7\" rel=\"nofollow noopener\" target=\"_blank\">1b<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">21<\/a>). Improvements were significant in all four major diseases (+19.9\u221225.0%, all P\u2009&lt;\u20090.001; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3a<\/a> and Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">17<\/a>\u2212<a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">20<\/a>). PCPs had disparate performance across skin tones as 4.6% (\u03b2\u2009=\u20090.046, 95% CI: 0.013\u22120.078, P\u2009=\u20090.069; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3c<\/a>), which was reduced to 2.9% (\u03b2\u2009=\u20090.029, 95% CI: \u22120.018 to 0.076, P\u2009=\u20090.248; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>) after AI assistance. Similar to the general public, no interactive effects of different XAI methods were found (P\u2009&gt;\u20090.05 for all conditions; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">22<\/a>), showing that the reduced disparities still resulted from CDANN.<\/p>\n<p>Fig. 3: Overall diagnostic performance with AI assistance for PCPs.<img decoding=\"async\" aria-describedby=\"figure-3-desc\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/08\/41591_2026_4553_Fig3_HTML.png\" alt=\"Fig. 3: Overall diagnostic performance with AI assistance for PCPs.\" loading=\"lazy\" width=\"685\" height=\"394\"\/><\/p>\n<p>This figure shows similar analysis results as in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig2\" rel=\"nofollow noopener\" target=\"_blank\">2<\/a> using the Human-Frst paradigm. a, Accuracy of PCPs (n\u2009=\u200996) in identifying atopic dermatitis, pityriasis rosea, Lyme disease and CTCL as well as the overall performance on four main diseases, with results presented for two rounds of testing. b, Diagnostic accuracy changes between Rd. 1 and Rd. 2 across different AI explanations (Basic (n\u2009=\u200924), GradCAM (n\u2009=\u200925), CBIR (n\u2009=\u200921) and LLM (n\u2009=\u200926), same as e and f). c, Diagnostic accuracy when patients are light-skinned (Fitzpatrick labels 1\u22124 on a scale of 6) or dark-skinned (Fitzpatrick labels 5\u22126) across the two rounds of testing. d, Overall accuracy (n\u2009=\u200996) changes between Rd. 1 and Rd. 2 under AI-right and AI-wrong conditions. e,f, Diagnostic accuracy changes between Rd. 1 and Rd. 2 of different AI explanations under AI-right (e) and AI-wrong (f) predictions. Each gray transparent scatter represents one measured sample. All presented bars or data points of a plot in this figure indicate the mean values of metrics measured. All error bars indicate the standard error of the mean. P values were calculated with post hoc marginal comparison after an initial analysis with a linear mixed model. ***P\u2009&lt;\u20090.001, **P\u2009&lt;\u20090.01, *P&lt;\u20090.05. NS, not significant; Rd., round.<\/p>\n<p><a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM5\" rel=\"nofollow noopener\" target=\"_blank\">Source data<\/a><\/p>\n<p>PCPs are resilient to AI deference when AI is wrong<\/p>\n<p>In contrast to the general public, incorrect AI predictions had minimal impact on the final decisions of PCPs across all XAI methods (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3f<\/a> and Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">2d\u2212f<\/a>; \u03b2\u2009=\u20090\u22120.021, P\u2009=\u20090.328\u22121.000). This suggests that PCPs relied on their own expertise and training rather than erroneous AI guidance. To control the effect of task difficulty between study 1 and study 2, we further compared the results of PCPs against the results of medical students on the same task. Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">5<\/a> shows that medical students relied on AI more than PCPs, regardless of AI correctness. We term participants who answered correctly only when AI was correct\u2014and, therefore, answered incorrectly when AI was wrong\u2014as \u2018deferential participants\u2019. We observed that the proportion of deferential participants was higher among medical students than PCPs (linear mixed model on the proportion of deferential participants, with medical role and medical expertise in skin as the main factor, controlling other confounders: main effect of medical role: \u03b2\u2009=\u20090.067, 95% CI: 0.004\u22120.130, P\u2009=\u20090.037; main effect of skin expertise: \u03b2\u2009=\u20090.073, 95% CI: 0.022\u22120.124, P\u2009=\u20090.005; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">25<\/a>). These results indicate that higher expertise levels are associated with more careful AI adoption, which is supported by previous work<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 10\" title=\"Tschandl, P. et al. Human&#x2013;computer collaboration for skin cancer recognition. Nat. Med. 26, 1229&#x2013;1234 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#ref-CR10\" id=\"ref-link-section-d67943945e1657\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>.<\/p>\n<p>LLM explanations do not aid in accuracy but in confidence calibration<\/p>\n<p>Interestingly, for PCPs, LLM explanations were the least helpful method (+17.7%) and were 8.1% lower than the best improvement from the basic explanations (\u03b2\u2009=\u2009\u22120.081, 95% CI: \u22120.207 to 0.044, P\u2009=\u20090.268; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig3\" rel=\"nofollow noopener\" target=\"_blank\">3b<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">16<\/a>). This finding was consistent across explanation quality (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig8\" rel=\"nofollow noopener\" target=\"_blank\">2e,g<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">27<\/a>) and top-3 accuracy (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig10\" rel=\"nofollow noopener\" target=\"_blank\">4c\u2212f<\/a>). These are opposite to the results of LLM\u02bcs best improvement for the general public in study 1. Medical student data confirmed that the task was not biased toward certain explanations (Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig11\" rel=\"nofollow noopener\" target=\"_blank\">5c<\/a>), implying that expertise drove interactions: the general public overrelies on LLMs, whereas PCPs are more resilient to incorrect LLM suggestions.<\/p>\n<p>However, LLM explanations did help improve the alignment between PCP participants\u02bc confidence and accuracy (correlation r\u2009=\u20090.494, P\u2009=\u20090.010; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">3d<\/a>; similar findings in top-3 performance; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">6d<\/a>) over No AI (correlation r\u2009=\u20090.084, P\u2009=\u20090.415, two-sided Fisher\u02bcs r-to-z test comparing the two correlations: z\u2009=\u20092.368, P\u2009=\u20090.018; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">3d<\/a>). This calibration benefit held even with incorrect AI predictions, where non-LLM explanations impaired alignment (Fisher\u02bcs r-to-z test z\u2009=\u20092.572, P\u2009=\u20090.010; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig9\" rel=\"nofollow noopener\" target=\"_blank\">3f<\/a>; similar in top-3 performance, P\u2009=\u20090.059; Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">6f<\/a>). PCP caution toward LLMs likely enforces cognitive engagement, thus enhancing diagnostic accuracy\u2212confidence calibration regardless of correctness<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 34\" title=\"Li, J., Yang, Y., Liao, Q. V., Zhang, J. &amp; Lee, Y.-C. As confidence aligns: understanding the effect of AI confidence on human self-confidence in human-AI decision making. In Proc. 2025 CHI Conference on Human Factors in Computing Systems (eds Evers, V. et al.) 1&#x2013;16 &#010;                https:\/\/doi.org\/10.1145\/3706598.3713336&#010;                &#010;               (Association for Computing Machinery, 2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#ref-CR34\" id=\"ref-link-section-d67943945e1754\" rel=\"nofollow noopener\" target=\"_blank\">34<\/a>,<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 35\" title=\"Said, A. On explaining recommendations with Large Language Models: a review. Front. Big Data 7, 1505284 (2025).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#ref-CR35\" id=\"ref-link-section-d67943945e1757\" rel=\"nofollow noopener\" target=\"_blank\">35<\/a>.<\/p>\n<p>Overall, in contrast to the general public in study 1, PCPs maintained their performance under incorrect AI predictions across all XAI methods, including LLM.<\/p>\n<p>Higher AI deference correlates with lower initial performance<\/p>\n<p>We inspected the relationship between participants\u02bc deference toward AI suggestions and initial performance<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 30\" title=\"Gaube, S. et al. Do as AI say: susceptibility in deployment of clinical decision-aids. NPJ Digit. Med. 4, 31 (2021).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#ref-CR30\" id=\"ref-link-section-d67943945e1773\" rel=\"nofollow noopener\" target=\"_blank\">30<\/a>. Among participants\u02bc final decisions that are correct (ranging from zero to 12, 12 images total), we visualize the number of images with AI suggestions that are correct (up to 10, shown in blue) and incorrect (up to two, shown in red) for each participant (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4a<\/a>). As mentioned above, \u2018deferential participants\u2019 are those who got correct results only when AI was correct and always got incorrect results when AI was wrong (that is, a blue bar in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4a,d<\/a>). Participants who got at least one correct outcome even when AI was incorrect are considered as \u2018non-deferential participants\u2019 (that is, a red bar on top of the blue bar in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4a,d<\/a>).<\/p>\n<p>Fig. 4: Deference patterns toward AI in human\u2212AI collaboration.<img decoding=\"async\" aria-describedby=\"figure-4-desc\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/08\/41591_2026_4553_Fig4_HTML.png\" alt=\"Fig. 4: Deference patterns toward AI in human&#x2212;AI collaboration.\" loading=\"lazy\" width=\"685\" height=\"363\"\/><\/p>\n<p>a,b, Performance comparison between deferential (fully compliant with AI suggestions) and non-deferential general populations. a, Distribution of correct responses with AI assistance. The blue bars represent instances where both AI\u02bcs predictions and participants\u02bc predictions were correct; the red overlays indicate instances where participants\u2019 responses were correct while AI predictions were incorrect. b, Accuracy trajectories of deferential users (n\u2009=\u2009142) and non-deferential users (n\u2009=\u2009193) found from a across two decision rounds. c, Proportion of participants who defer to the output of each type of XAI (Basic (n\u2009=\u200982), GradCAM (n\u2009=\u200972), CBIR (n\u2009=\u200991) and LLM (n\u2009=\u200990)). d\u2212f, Equivalent results for PCP participants as in a\u2212c (in e, deferential PCPs n\u2009=\u200979, non-deferential PCPs n\u2009=\u200917; in f, Basic (n\u2009=\u200924), GradCAM (n\u2009=\u200925), CBIR (n\u2009=\u200921) and LLM (n\u2009=\u200926)). All presented bars or data points of a plot in this figure indicate the mean values of metrics measured. All error bars indicate the standard error of the mean. P values were calculated with post hoc marginal comparison after an initial analysis with a linear mixed model. ***P\u2009&lt;\u20090.001, **P\u2009&lt;\u20090.01, *P\u2009&lt;\u20090.05. NS, not significant; Rd., round.<\/p>\n<p><a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM6\" rel=\"nofollow noopener\" target=\"_blank\">Source data<\/a><\/p>\n<p>Although the general public had similar performance after AI assistance, deferential participants had significantly lower initial diagnostic accuracy (65.8%) compared to non-deferential participants (72.5%) before receiving AI suggestions (\u03b2\u2009=\u20090.082, 95% CI: 0.074\u22120.111, P\u2009&lt;\u20090.001, Cohen\u2019s d\u2009=\u20090.462; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4b<\/a>; linear mixed model with deferential group as the main factor, controlling the same confounders as study 1 analysis) (Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">28<\/a>). Deferential PCPs also had a significantly lower performance in the initial round (\u03b2\u2009=\u20090.194, 95% CI: 0.079\u22120.310, P\u2009&lt;\u20090.001, Cohen\u2019s d\u2009=\u20091.578; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4e<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">29<\/a>), which could result from a lower level of critical thinking (P\u2009=\u20090.028, measured by critical thinking questionnaire<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 36\" title=\"Frederick, S. Cognitive reflection and decision making. J. Econ. Perspect. 19, 25&#x2013;42 (2005).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#ref-CR36\" id=\"ref-link-section-d67943945e1935\" rel=\"nofollow noopener\" target=\"_blank\">36<\/a>; see details in <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"section anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Sec16\" rel=\"nofollow noopener\" target=\"_blank\">Methods<\/a>).<\/p>\n<p>LLM explanations led to the largest proportion of fully deferential participants in the general public (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4c<\/a>), although no significance between LLM and others was observed (P\u2009=\u20090.169\u22120.433; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">30<\/a>). By contrast, LLM resulted in the lowest proportion of deferential PCPs (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4f<\/a>; P\u2009=\u20090.268\u22120.587; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">31<\/a>). This is also aligned with our findings of the capability of experts in maintaining resilience against the misdirection of wrong semantic AI explanations<a data-track=\"click\" data-track-action=\"reference anchor\" data-track-label=\"link\" data-test=\"citation-ref\" aria-label=\"Reference 10\" title=\"Tschandl, P. et al. Human&#x2013;computer collaboration for skin cancer recognition. Nat. Med. 26, 1229&#x2013;1234 (2020).\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#ref-CR10\" id=\"ref-link-section-d67943945e1964\" rel=\"nofollow noopener\" target=\"_blank\">10<\/a>.<\/p>\n<p>Putting AI before human decisions amplifies deference across expertise levels<\/p>\n<p>In addition to XAI methods, a practical design factor for human\u2212AI collaboration systems is the decision-making order, either Human-First or AI-First paradigms. Both the general public and PCPs had significantly better performance in the first round with AI-First (all P\u2009&lt;\u20090.001; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5a,d<\/a>), which is not surprising due to superior AI performance. In the second round, after humans received the same amount of information, no difference was observed between Human-First and AI-First in either study (round 2, general public: \u03b2\u2009=\u20090.005, 95% CI: \u22120.017 to 0.026, P\u2009=\u20090.650; PCPs: \u03b2\u2009=\u2009\u22120.020, 95% CI: \u22120.091 to 0.050, P\u2009=\u20090.572; linear mixed models on accuracy with human\u2212AI collaboration paradigm, decision round and their interaction as the main factors, controlling decision-making time and other confounders; see Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">32<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">34<\/a> for details). This indicates that the decision order may not influence the final performance. To exclude the influence where participants would be biased in the second round due to the prior exposure of the disease image, we compared AI-assisted performance in AI-First round 1 versus performance in Human-First round 2 and found no differences (P\u2009&gt;\u20090.05 for both general public and PCPs; Supplementary Tables <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">38<\/a> and <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">39<\/a>), indicating that the diagnosis strategies were not driven by the carryover effect.<\/p>\n<p>Fig. 5: Influence of human\u2212AI collaboration paradigm on diagnostic performance.<img decoding=\"async\" aria-describedby=\"figure-5-desc\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/08\/41591_2026_4553_Fig5_HTML.png\" alt=\"Fig. 5: Influence of human&#x2212;AI collaboration paradigm on diagnostic performance.\" loading=\"lazy\" width=\"685\" height=\"358\"\/><\/p>\n<p>a\u2212d, Analysis for the general public. a, Comparison between AI-First (n\u2009=\u2009288) and Human-First (n\u2009=\u2009335) human\u2212AI collaboration (HAI) systems on diagnostic accuracy across two rounds. b, Accuracy trajectories across two decision rounds for deferential and non-deferential people in the AI-First paradigm. c, Comparison of deferential user (participants who defer to the output of each type of XAI) proportion between two HAI paradigms across four XAI groups (AI-First: Basic (n\u2009=\u200962), GradCAM (n\u2009=\u200974), CBIR (n\u2009=\u200976) and LLM (n\u2009=\u200976); Human-First: Basic (n\u2009=\u200982), GradCAM (n\u2009=\u200972), CBIR (n\u2009=\u200991) and LLM (n\u2009=\u200990)). d\u2212f, Equivalent results for PCP participants (AI-First PCPs n\u2009=\u200957, among which Basic (n\u2009=\u200914), GradCAM (n\u2009=\u200918), CBIR (n\u2009=\u20098) and LLM (n\u2009=\u200917); Human-First PCPs n\u2009=\u200996, among which Basic (n\u2009=\u200924), GradCAM (n\u2009=\u200925), CBIR (n\u2009=\u200921) and LLM (n\u2009=\u200926)) as a\u2212c. All presented bars or data points of a plot in this figure indicate the mean values of metrics measured. All error bars indicate the standard error of the mean. P values were calculated with post hoc marginal comparison after an initial analysis with a linear mixed model. ***P\u2009&lt;\u20090.001, **P\u2009&lt;\u20090.01, *P\u2009&lt;\u20090.05. NS, not significant; Rd., round.<\/p>\n<p><a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM7\" rel=\"nofollow noopener\" target=\"_blank\">Source data<\/a><\/p>\n<p>With the AI-First paradigm, non-deferential participants still had better performance, especially after reviewing the examples again without AI (general public: \u03b2\u2009=\u20090.031, 95% CI: 0.000\u22120.062, P\u2009=\u20090.049; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">33<\/a>; PCPs: \u03b2\u2009=\u20090.218, 95% CI: 0.009\u22120.426, P\u2009=\u20090.041; Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">35<\/a>). This is similar to the results in the Human-First paradigm in Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig4\" rel=\"nofollow noopener\" target=\"_blank\">4<\/a>. By contrast, putting AI suggestions ahead increased the proportion of deferential participants in most cases (Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5c,f<\/a>). For the general public, the proportion was increased across all XAI methods (average \u0394\u2009=\u2009+8.4%, although no significance was observed after controlling all confounders, P\u2009=\u20090.067\u22120.170; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5c<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">36<\/a>). For PCPs, the largest deference increase was observed from LLM explanations (\u0394\u2009=\u2009+19.0%; Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig5\" rel=\"nofollow noopener\" target=\"_blank\">5f<\/a>, Extended Data Fig. <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig12\" rel=\"nofollow noopener\" target=\"_blank\">6m<\/a> and Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">37<\/a>). This indicates that putting AI ahead may lead to stronger anchoring bias. It also suggests that, although PCPs showed resistance against AI\u02bcs mislead in the Human-First paradigm, providing LLM-based explanations ahead of human choices can still cause more bias than other XAI methods and introduce risks of overreliance, even for PCPs.<\/p>\n<p>Human\u2212AI collaboration to combine each side\u02bcs strength<\/p>\n<p>Although AI deference risks misleading participants when AI makes mistakes, deference may lead some humans to improved performance. Figure <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"figure anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#Fig6\" rel=\"nofollow noopener\" target=\"_blank\">6<\/a> visualizes cases where either humans or AI routinely outperform each other. We found that AI tends to outperform humans in cases where the presentation of the disease is subtle but struggles with atypical symptoms or unexpected features in the image (see Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">40<\/a> for example information). These qualitative examples provide some initial directions for future work in understanding the complementary strengths of humans and AI in dermatological diagnosis.<\/p>\n<p>Fig. 6: Case study characterizing when AI or humans did better.<img decoding=\"async\" aria-describedby=\"figure-6-desc\" src=\"https:\/\/www.newsbeep.com\/ca\/wp-content\/uploads\/2026\/08\/41591_2026_4553_Fig6_HTML.png\" alt=\"Fig. 6: Case study characterizing when AI or humans did better.\" loading=\"lazy\" width=\"685\" height=\"484\"\/><\/p>\n<p>a, Accuracy proportion of images in study 1 for the general public, categorized by whether the AI prediction was correct or incorrect. The blue bars represent the proportion of correct human participants\u02bc decisions; the red overlays indicate their incorrect decisions. Representative images are labeled as \u2018AI did better than human\u2019 (AI-right images with the highest human error rate), and \u2018Human did better than AI\u2019 (AI-wrong images with the highest human accuracy) is shown to the right. b, Same as a for study 2 for PCPs. Detailed examples can be found in Supplementary Table <a data-track=\"click\" data-track-label=\"link\" data-track-action=\"supplementary material anchor\" href=\"http:\/\/www.nature.com\/articles\/s41591-026-04553-w#MOESM1\" rel=\"nofollow noopener\" target=\"_blank\">40<\/a>. pred, prediction.<\/p>\n","protected":false},"excerpt":{"rendered":"Study design We designed two complementary large-scale experiments to evaluate human\u2212AI collaborative diagnostic performance across expertise levels (Fig.&hellip;\n","protected":false},"author":2,"featured_media":846583,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[34],"tags":[3875,49,48,3877,3673,84,392,2724,796,6918,6919,6920,294744,294745],"class_list":["post-846582","post","type-post","status-publish","format-standard","has-post-thumbnail","category-healthcare","tag-biomedicine","tag-ca","tag-canada","tag-cancer-research","tag-general","tag-health","tag-healthcare","tag-infectious-diseases","tag-machine-learning","tag-metabolic-diseases","tag-molecular-medicine","tag-neurosciences","tag-predictive-medicine","tag-skin-manifestations"],"_links":{"self":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/846582","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/comments?post=846582"}],"version-history":[{"count":0,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/posts\/846582\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media\/846583"}],"wp:attachment":[{"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/media?parent=846582"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/categories?post=846582"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.newsbeep.com\/ca\/wp-json\/wp\/v2\/tags?post=846582"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}