A Human-Centered Validation of the Explainability-Performance Coefficient
A model-agnostic metric, the EPC score, extending the Explainability-Performance Coefficient, is proposed to quantify explanation quality by balancing feature selection sparsity and preserved model performance. Empirical validation across tabular, text, and image modalities shows the EPC score uncovers operational dependencies among network activations, data dimensionality, and explainer performance. Validation against independent human-based explanations proves higher EPC scores strongly align with human lexical sentiment judgments and spatial visual annotations.
Researchers propose the EPC score, a model-agnostic extension of the Explainability-Performance Coefficient, to objectively evaluate explanation fidelity in deep learning models. The metric balances sparsity and performance, and is validated across multiple data modalities. Human-centered validation demonstrates strong alignment between higher EPC scores and human judgments, including lexical sentiment and spatial annotations.
The EPC score explicitly models the trade-off between feature selection sparsity and preserved model performance, providing a unified metric for explanation quality. Its model-agnostic nature allows application across tabular, text, and image modalities, revealing dependencies between network activations, data dimensionality, and explainer performance.
This work addresses a critical gap in trustworthy AI for high-risk domains by providing an objective, human-aligned metric for XAI evaluation. Adoption could standardize explanation quality assessment, facilitating regulatory compliance and user trust in sectors like healthcare, finance, and autonomous systems.
The EPC score offers a quantifiable way to measure and compare explainability across models, which can reduce risk in deploying AI in regulated industries, improve model auditing processes, and enhance end-user confidence in AI-driven decisions.
Next signals include potential integration of the EPC score into XAI toolkits and benchmarks, further validation on larger-scale human studies, and exploration of its use in model selection and debugging workflows. Regulatory bodies may reference such metrics in upcoming AI governance frameworks.