Event date · · GPT-4

Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support

Healthcare Ai
FACT STATEMENT

A retrospective study of 99 ED revisit diagnosis pairs found GPT-4 rated 94% as warranting follow-up, 4.4-13.3 times more than clinicians, with poor correlation to clinician raters. An LLM-populated knowledge graph algorithm (KGA) was created to screen potentially concerning pairs.

What happened

Researchers conducted an exploratory retrospective study of emergency department revisits within 1-14 days. Clinicians and GPT-4 assessed whether diagnosis pairs warranted further review. GPT-4 over-flagged nearly all pairs (94%) compared to clinicians, showing poor correlation. An algorithm using an LLM-populated knowledge graph was developed to automatically screen concerning pairs.

Technical significance

GPT-4's high false-positive rate suggests minimal prompt engineering and lack of calibration for clinical triage tasks. The KGA approach attempts to combine LLM knowledge with structured graph reasoning to improve screening specificity.

Industry impact

This highlights the gap between general LLM capabilities and specialized clinical decision support. Over-flagging could increase clinician burden rather than reduce it, underscoring the need for domain-specific fine-tuning and evaluation.

Decision value

If KGA reduces false positives while maintaining sensitivity, it could lower chart review costs and improve quality assurance efficiency in health systems. However, current GPT-4 behavior would likely be counterproductive without refinement.

What to watch

Next signals include publication of KGA performance metrics, comparison with rule-based screening, and potential prospective validation in hospital settings. Watch for improved prompting or fine-tuned medical LLMs addressing calibration.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.