Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods
A paper published on arXiv on 2026-09-04 introduces behavioral correctness assumptions as a framework for evaluating reference-based automatic evaluation methods. It defines a taxonomy of correctness-preserving and correctness-altering assumptions and operationalizes them through controlled response transformations. The study evaluates diverse lexical, character-level, semantic, LLM-based, and hybrid evaluators, analyzing assumption-level behavior, stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility. Findings show no evaluator satisfies all proposed correctness assumptions, and evaluators with similar aggregate performance can have substantially different behavioral profiles.
The paper proposes a complementary framework to existing meta-evaluation methods by focusing on behavioral correctness assumptions rather than aggregate agreement with human judgments. It systematically tests various automatic evaluation methods under controlled transformations to reveal their behavioral profiles. The results indicate that current evaluators exhibit distinct trade-offs and that aggregate performance metrics can mask important behavioral differences, providing diagnostic information for evaluator selection and development.
The framework operationalizes correctness assumptions through controlled response transformations, enabling fine-grained analysis of evaluator behavior. The study covers a wide range of evaluator types, including lexical, character-level, semantic, LLM-based, and hybrid methods. Key technical dimensions analyzed include stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility. The finding that no evaluator satisfies all assumptions suggests inherent limitations in current evaluation paradigms and highlights the need for assumption-aware evaluator design.
For AI practitioners, this research implies that relying solely on aggregate benchmark scores may be insufficient for selecting evaluation methods. The behavioral profiles can guide the choice of evaluator based on specific application requirements, such as robustness to paraphrasing or sensitivity to semantic changes. This could influence tooling and best practices in NLG system development, encouraging more nuanced evaluation pipelines.
The framework provides a method for more reliable evaluator selection, potentially reducing risks in deploying NLG systems where evaluation quality directly impacts product performance. Companies developing evaluation tools or relying on automatic metrics can use behavioral profiles to differentiate offerings or improve internal evaluation pipelines. The research may also inform procurement decisions for evaluation infrastructure.
Future work may extend the taxonomy of correctness assumptions and apply the framework to emerging evaluator types, including more advanced LLM-based and hybrid systems. The diagnostic nature of the framework could lead to standardized behavioral testing suites for automatic evaluation methods. Observables include adoption of the framework in meta-evaluation benchmarks and development of evaluators that explicitly target a broader set of correctness assumptions.