Hoot, N. R. & Aronsky, D. Systematic review of emergency department crowding: causes, effects, and solutions. Ann. Emerg. Med. 52, 126–136 (2008).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Morley, C., Unwin, M., Peterson, G. M., Stankovich, J. & Kinsman, L. Emergency department crowding: a systematic review of causes, consequences and solutions. PLoS ONE 13, e0203316 (2018).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Voaklander, B. et al. Interventions to improve consultations in the emergency department: a systematic review. Acad. Emerg. Med. 29, 1475–1495 (2022).

Article 
PubMed 

Google Scholar
 

Brick, C. et al. The impact of consultation on length of stay in tertiary care emergency departments. Emerg. Med. J. 31, 134–138 (2014).

Article 
PubMed 

Google Scholar
 

Piliuk, K. & Tomforde, S. Artificial intelligence in emergency medicine: a systematic literature review. Int. J. Med. Inform. 180, 105274 (2023).

Article 
PubMed 

Google Scholar
 

Grant, K., McParland, A., Mehta, S. & Ackery, A. D. Artificial intelligence in emergency medicine: surmountable barriers with revolutionary potential. Ann. Emerg. Med. 75, 721–726 (2020).

Article 
PubMed 

Google Scholar
 

Nagendran, M. et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies. BMJ 368, m689 (2020).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Plana, D. et al. Randomized clinical trials of machine learning interventions in health care: a systematic review. JAMA Netw. Open 5, e2233946 (2022).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Han, T. et al. Randomised controlled trials evaluating artificial intelligence in clinical practice: a scoping review. Lancet Digit. Health 6, e367–e373 (2024).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Oikonomidi, T. et al. Applications of artificial intelligence for real-world evidence generation: a protocol for a living scoping review. BMJ Open 16, e109725 (2026).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Kelly, C. J., Karthikesalingam, A., Suleyman, M., Corrado, G. & King, D. Key challenges for delivering clinical impact with artificial intelligence. BMC Med. 17, 195 (2019).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Wiens, J. et al. Do no harm: a roadmap for responsible machine learning for health care. Nat. Med. 25, 1337–1340 (2019).

Article 
CAS 
PubMed 

Google Scholar
 

Vasey, B. et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 28, 924–933 (2022).

Article 
CAS 
PubMed 

Google Scholar
 

Thirunavukarasu, A. J. et al. Large language models in medicine. Nat. Med. 29, 1930–1940 (2023).

Article 
CAS 
PubMed 

Google Scholar
 

Singhal, K. et al. Large language models encode clinical knowledge. Nature 620, 172–180 (2023).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Jiang, L. Y. et al. Health system-scale language models are all-purpose prediction engines. Nature 619, 357–362 (2023).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Williams, C. Y. K., Miao, B. Y., Kornblith, A. E. & Butte, A. J. Evaluating the use of large language models to provide clinical recommendations in the emergency department. Nat. Commun. 15, 8236 (2024).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Hager, P. et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nat. Med. 30, 2613–2622 (2024).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Preiksaitis, C. & Rose, C. The role of large language models in transforming emergency medicine: scoping review. JMIR Med. Inform. 12, e53787 (2024).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Gilboy, N., Tanabe, T., Travers, D. & Rosenau, A. M.Emergency Severity Index (ESI): A Triage Tool for Emergency Department Care, Version 4. AHRQ Publication No. 12-0014 (AHRQ, 2011).

Greenhalgh, T. et al. Beyond adoption: a new framework for theorizing and evaluating nonadoption, abandonment, and challenges to the scale-up, spread, and sustainability of health and care technologies. J. Med. Internet Res. 19, e367 (2017).

Article 
PubMed 
PubMed Central 

Google Scholar
 

R Core Team R: A Language and Environment for Statistical Computing (R Core Team, 2024).

Ancker, J. S. et al. Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system. BMC Med. Inform. Decis. Mak. 17, 36 (2017).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Kawamoto, K., Houlihan, C. A., Balas, E. A. & Lobach, D. F. Improving clinical practice using clinical decision support systems: a systematic review of trials to identify features critical to success. BMJ 330, 765 (2005).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Roshanov, P. S. et al. Features of effective computerised clinical decision support systems: meta-regression of 162 randomised trials. BMJ 346, f657 (2013).

Article 
PubMed 

Google Scholar
 

Proctor, E. K. et al. Outcomes for implementation research: conceptual distinctions, measurement challenges, and research agenda. Adm. Policy Ment. Health 38, 65–76 (2011).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Angrist, J. D., Imbens, G. W. & Rubin, D. B. Identification of causal effects using instrumental variables. J. Am. Stat. Assoc. 91, 444–455 (1996).

Article 

Google Scholar
 

Kwong, J. C. C., Nickel, G. C., Wang, S. C. Y. & Kvedar, J. C. Integrating artificial intelligence into healthcare systems: more than just the algorithm. npj Digit. Med. 7, 52 (2024).

Article 
PubMed 
PubMed Central 

Google Scholar
 

McCambridge, J., Witton, J. & Elbourne, D. R. Systematic review of the Hawthorne effect: new concepts are needed to study research participation effects. J. Clin. Epidemiol. 67, 267–277 (2014).

Article 
PubMed 

Google Scholar
 

Geskey, J. M., Geeting, G., West, C. & Hollenbeak, C. S. Improved physician consult response times in an academic emergency department after implementation of an institutional guideline. J. Emerg. Med. 44, 999–1006 (2013).

Article 
PubMed 

Google Scholar
 

Soong, C., High, S., Morgan, M. W. & Ovens, H. A novel approach to improving emergency department consultant response times. BMJ Qual. Saf. 22, 299–305 (2013).

Article 
PubMed 

Google Scholar
 

VanderWeele, T. J. & Ding, P. Sensitivity analysis in observational research: introducing the E-value. Ann. Intern. Med. 167, 268–274 (2017).

Article 
PubMed 

Google Scholar
 

Donabedian, A. The quality of care: how can it be assessed?. JAMA 260, 1743–1748 (1988).

Article 
CAS 
PubMed 

Google Scholar
 

Mant, J. Process versus outcome indicators in the assessment of quality of health care. Int. J. Qual. Health Care 13, 475–480 (2001).

Article 
CAS 
PubMed 

Google Scholar
 

Torgerson, D. J. Contamination in trials: is cluster randomisation the answer? BMJ 322, 355–357 (2001).

Article 
CAS 
PubMed 
PubMed Central 

Google Scholar
 

Hemming, K., Haines, T. P., Chilton, P. J., Girling, A. J. & Lilford, R. J. The stepped wedge cluster randomised trial: rationale, design, analysis, and reporting. BMJ 350, h391 (2015).

Article 
CAS 
PubMed 

Google Scholar
 

Lewis P. et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. In Proc. 34th International Conference on Neural Information Processing Systems (NIPS ’20) (eds Larochelle, H. et al.) 9459–9474 (Curran, 2020).

Topol, E. J. High-performance medicine: the convergence of human and artificial intelligence. Nat. Med. 25, 44–56 (2019).

Article 
CAS 
PubMed 

Google Scholar
 

Ethics and Governance of Artificial Intelligence for Health: WHO Guidance (WHO, 2021).

Shin, S., Lee, S. H., Kim, D. H. & Kim, K. H. The impact of the improvement in internal medicine consultation process on emergency department length of stay. Am. J. Emerg. Med. 36, 620–624 (2018).

Article 
PubMed 

Google Scholar
 

Ravi, A., Shochat, G., Wang, R. C. & Khanna, R. Improvements to emergency department length of stay and user satisfaction after implementation of an integrated consult order. J. Am. Coll. Emerg. Physicians Open 4, e12922 (2023).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Austin, P. C. Balance diagnostics for comparing the distribution of baseline covariates between treatment groups in propensity-score matched samples. Stat. Med. 28, 3083–3107 (2009).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Greenland, S. An introduction to instrumental variables for epidemiologists. Int. J. Epidemiol. 29, 722–729 (2000).

Article 
CAS 
PubMed 

Google Scholar
 

Staiger, D. & Stock, J. H. Instrumental variables regression with weak instruments. Econometrica 65, 557–586 (1997).

Article 

Google Scholar
 

Hodges, J. L. Jr & Lehmann, E. L. Estimates of location based on rank tests. Ann. Math. Stat. 34, 598–611 (1963).

Article 

Google Scholar
 

Meunier, P. Y., Raynaud, C., Guimaraes, E., Gueyffier, F. & Letrilliart, L. Barriers and facilitators to the use of clinical decision support systems in primary care: a mixed-methods systematic review. Ann. Fam. Med. 21, 57–69 (2023).

Article 
PubMed 
PubMed Central 

Google Scholar
 

Python v.3.13 (Python Software Foundation, 2024).

llironlibo. llironlibo/SHAKED-analysis: v1.1.0. Zenodo https://doi.org/10.5281/zenodo.20736931 (2026).