Event date · · arXiv

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

FACT STATEMENT

A study published on arXiv on 2026-08-10 introduces a linguistically grounded annotation schema with 10 perceptual dimensions to evaluate automated TTS evaluators. The benchmark includes 860 utterances annotated by trained linguists. Four MOS predictors and four Audio-LLM judges were tested. MOS predictors primarily reflect acoustic signal quality, while Audio-LLM judges show prompt-dependent detection that does not generalize across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. The dataset, schema, and code are publicly released.

What happened

Researchers have deconstructed the concept of 'naturalness' in text-to-speech evaluation into 10 distinct perceptual dimensions, creating a meta-evaluation benchmark of 860 utterances annotated by linguists. Testing four MOS predictors and four Audio-LLM judges revealed that MOS predictors focus narrowly on acoustic quality, while Audio-LLM judges exhibit selective, prompt-dependent performance. Neither approach reliably detects the full range of linguistically structured speech errors, highlighting a gap in current automated TTS evaluation methods.

Technical significance

The study demonstrates that current automated TTS evaluators, including MOS predictors and Audio-LLM judges, fail to capture linguistically grounded perceptual dimensions beyond basic acoustic quality. MOS predictors collapse onto signal-level features, while Audio-LLM judges' performance varies with prompting and does not generalize across all dimensions. This suggests that existing models lack the fine-grained linguistic understanding required for comprehensive speech quality assessment.

Industry impact

The findings indicate that the TTS industry's reliance on automated evaluation metrics may overlook critical aspects of speech naturalness, potentially leading to systems that sound superficially good but contain subtle linguistic errors. This could impact user trust and adoption in applications like virtual assistants, audiobooks, and accessibility tools. The public release of the benchmark may drive development of more robust evaluation methods.

Decision value

For companies developing TTS systems, adopting more comprehensive evaluation methods could lead to higher-quality speech synthesis, improving user experience and competitive differentiation. The benchmark provides a tool for identifying and correcting specific linguistic errors, potentially reducing the need for costly human evaluation. It may also open opportunities for new evaluation-as-a-service offerings.

What to watch

Future work may focus on developing new automated evaluators that incorporate linguistic dimensions, possibly through multi-task learning or fine-tuning on the released dataset. We may see increased collaboration between linguists and AI researchers to refine evaluation schemas. The benchmark could become a standard for TTS evaluation, influencing both academic research and industry practices.

DECISION BRIEF

Turn the evidence into a decision.

See how AIGC.NEWS separates verified change, judgment, and the next signal to watch.