FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection
The study FOUND-AF evaluates nine publicly available ECG foundation models from five families (HuBERT-ECG, CLEF, ST-MEM, ECG-JEPA, ECGFounder) on four heterogeneous ECG datasets (AFDB, CinC2017, CPSC2021, LTAFDB) using a unified, leakage-controlled, deployment-oriented benchmarking framework. All models were used as frozen feature extractors with standardized preprocessing, model-native resampling, a fixed XGBoost classifier, and recording-level grouped cross-validation. Evaluation included classification metrics, ROC analysis, paired recording-level bootstrap comparisons with Holm correction, and embedding-space analysis.
Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia, associated with increased risks of stroke, heart failure, and mortality. ECG foundation models offer transferable representations for automated AF detection, but their relative effectiveness has been unclear due to inconsistent evaluation protocols. FOUND-AF provides a unified benchmarking framework to compare nine models across four datasets under identical conditions, revealing performance differences and guiding model selection for clinical deployment.
The framework enforces leakage-controlled, recording-level grouped cross-validation and uses frozen feature extractors with a fixed XGBoost classifier, ensuring fair comparison. Embedding-space analysis and paired bootstrap tests with Holm correction provide rigorous statistical evidence of model performance differences.
Standardized benchmarking of ECG foundation models can accelerate adoption of AI-assisted AF detection in clinical settings by identifying the most robust and generalizable models, reducing the risk of deploying underperforming systems.
Healthcare providers and AI developers can use the benchmark results to select optimal ECG foundation models for AF detection, potentially improving diagnostic accuracy, reducing costs, and enhancing patient outcomes.
Future work may extend the benchmark to additional datasets, real-world deployment scenarios, and other cardiac conditions. The framework could become a standard for evaluating ECG AI models, influencing regulatory approval and clinical integration.