MP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant
MP-Bench is introduced as the first benchmark specifically designed to objectively evaluate conversational speech systems as active participants within multi-party contexts. It assesses agent behavior along two primary dimensions: turn-taking awareness and response appropriateness, and incorporates comprehension-based question-answering tasks as a complementary evaluation. The benchmark was used to evaluate 12 voice agents.
Conversational voice agents have advanced significantly, but existing benchmarks largely overlook multi-party conversations. MP-Bench addresses this gap by evaluating agents on turn-taking awareness, response appropriateness, and comprehension in multi-party settings. Benchmarking 12 voice agents reveals that real-time voice agents face challenges in these complex scenarios.
The benchmark evaluates two key dimensions: turn-taking awareness and response appropriateness, plus comprehension-based QA. This suggests a focus on open turn-taking dynamics and contextual response generation, which are more complex than dyadic interactions. The evaluation of 12 agents indicates a need for improved handling of conversational complexity in multi-party settings.
The introduction of MP-Bench signals a shift in voice agent evaluation from dyadic to multi-party scenarios, reflecting real-world use cases like meetings and group interactions. This may drive development of more socially aware voice agents and influence industry standards for conversational AI.
MP-Bench provides a standardized way to assess voice agents for multi-party applications, potentially reducing risk for enterprises deploying such systems in collaborative environments. It may also serve as a differentiator for vendors targeting group communication tools.
Future work may expand MP-Bench to more diverse multi-party scenarios and languages, and incorporate additional metrics such as interruption handling and speaker diarization. The benchmark could become a standard for evaluating voice agents in group settings, prompting improvements in turn-taking models.