EDRAC: Benchmarking Arabic Dialect Reading Comprehension
EDRAC is introduced as the first large-scale benchmark for dialectal Arabic machine reading comprehension and generative QA, covering Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic. It contains 499 passages from naturally occurring spoken interactions and 4,977 QA pairs generated through a human-LLM collaborative pipeline. Benchmarking of Arabic-centric and multilingual LLMs reveals substantial gaps between semantic answer quality and dialectal fidelity.
Dialectal Arabic remains under-resourced compared to Modern Standard Arabic, especially for machine reading comprehension and question answering. EDRAC addresses this gap with a benchmark of 499 spoken-interaction passages and 4,977 QA pairs across five major dialects. Evaluation of Arabic-centric and multilingual LLMs shows significant discrepancies between semantic quality and dialectal fidelity, highlighting limitations of current metrics for dialectal Arabic generation.
The human-LLM collaborative pipeline combines iterative generation, LLM-as-a-judge evaluation, and human verification to create QA pairs. Lexical and semantic metrics are used for benchmarking, revealing that existing evaluation metrics are insufficient for capturing dialectal fidelity in generated answers.
The benchmark exposes a clear performance gap in dialectal Arabic NLP, indicating that current LLMs are not yet robust for real-world dialectal applications. This creates an opportunity for model developers to improve dialectal understanding and generation, particularly for Arabic-speaking markets.
Improved dialectal Arabic NLP could enable better customer service, content moderation, and information access for Arabic-speaking populations. Companies targeting Middle East and North Africa markets may benefit from models that handle dialectal variations effectively.
Future research may focus on developing better evaluation metrics for dialectal Arabic generation and improving LLM performance on dialectal MRC. Adoption of EDRAC as a standard benchmark could drive progress in under-resourced Arabic dialects.