MedHELM: A Comprehensive Evaluation Framework for Medical LLMs
In May 2025, Stanford and other institutions released MedHELM, a medical LLM evaluation framework validated by 29 clinicians, covering 5 categories, 22 subcategories, and 121 tasks, integrating 35 benchmarks (17 existing + 18 newly constructed). Evaluating 9 frontier LLMs found: reasoning models (DeepSeek R1 win rate 66%, o3-mini win rate 64%) performed best, but Claude 3.5 Sonnet achieved comparable performance at 40% lower computational cost. Models performed strongly on clinical note generation (0.73-0.85) and patient communication (0.78-0.83), but weakly on clinical decision support (0.56-0.72) and administrative workflows (0.53-0.63). The LLM-jury evaluation method showed agreement with clinician scores (ICC=0.47) exceeding inter-clinician agreement (ICC=0.43).
MedHELM is the most comprehensive medical LLM evaluation benchmark to date, revealing significant performance differences across models in real clinical tasks. Key findings: reasoning models lead overall, but Claude 3.5 Sonnet is more cost-effective; all models underperform on core tasks like clinical decision support, indicating current LLMs are not yet reliable for diagnostic assistance. The LLM-jury method offers a new approach for automated evaluation, with agreement even surpassing that among human experts. This framework provides a scientific basis for procurement, deployment, and regulation of medical AI.
MedHELM framework construction: (1) 29 clinicians used the Delphi method to determine the task taxonomy, covering 5 categories: clinical note generation, patient communication and education, medical research assistance, clinical decision support, and administration and workflow. (2) Integrated 35 benchmarks, 18 newly constructed, ensuring at least one benchmark per subcategory. (3) Evaluation method uses LLM-jury: multiple LLMs as reviewers score model outputs, taking weighted averages. Results show LLM-jury agreement with clinician scores (ICC=0.47) outperforms traditional automatic metrics (ROUGE-L 0.36, BERTScore-F1 0.44) and inter-clinician agreement (0.43). Model performance: reasoning models lead on complex reasoning tasks, but Claude 3.5 Sonnet achieves similar performance at lower cost on most tasks.
MedHELM provides a standardized evaluation tool for the medical AI industry, helping healthcare institutions and regulators select models suitable for clinical scenarios. Its cost-performance analysis directly guides procurement decisions: for hospitals with limited budgets, Claude 3.5 Sonnet may be the most cost-effective choice. The framework's openness (open source) will drive more medical AI evaluation research and accelerate industry standardization.
Recommend that medical IT procurement departments use the MedHELM framework to evaluate candidate LLMs, focusing on clinical decision support subcategory scores. For non-critical tasks (e.g., clinical note generation), deploy Claude 3.5 Sonnet to reduce costs. For high-risk scenarios like diagnostic assistance, current models are not yet mature; deployment should be postponed or used only as a supplementary reference.
Watch whether MedHELM is adopted by regulators like the FDA as a reference for approval. Observe the generalization ability of the LLM-jury method in more medical scenarios (e.g., imaging diagnosis). Low scores on clinical decision support suggest a need for specialized training or fine-tuning; future dedicated medical models for specific departments may emerge.