Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support
A retrospective study of 99 ED revisit diagnosis pairs found GPT-4 rated 94% as warranting follow-up, 4.4-13.3 times more than clinicians, with poor correlation to clinician raters. An LLM-populated knowledge graph algorithm (KGA) was created to screen potentially concerning pairs.
Researchers conducted an exploratory retrospective study of emergency department revisits within 1-14 days. Clinicians and GPT-4 assessed whether diagnosis pairs warranted further review. GPT-4 over-flagged nearly all pairs (94%) compared to clinicians, showing poor correlation. An algorithm using an LLM-populated knowledge graph was developed to automatically screen concerning pairs.
GPT-4's high false-positive rate suggests minimal prompt engineering and lack of calibration for clinical triage tasks. The KGA approach attempts to combine LLM knowledge with structured graph reasoning to improve screening specificity.
This highlights the gap between general LLM capabilities and specialized clinical decision support. Over-flagging could increase clinician burden rather than reduce it, underscoring the need for domain-specific fine-tuning and evaluation.
If KGA reduces false positives while maintaining sensitivity, it could lower chart review costs and improve quality assurance efficiency in health systems. However, current GPT-4 behavior would likely be counterproductive without refinement.
Next signals include publication of KGA performance metrics, comparison with rule-based screening, and potential prospective validation in hospital settings. Watch for improved prompting or fine-tuned medical LLMs addressing calibration.