FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
FriendBench is a benchmark for inferring whether two people are familiar or strangers from a 20-second clip of a dyadic ice-breaker conversation. It compares 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but the strongest models lean toward 'stranger' while humans stay balanced. Richer channels help both unequally, and only humans gain from visible behavior on top of speech.
FriendBench evaluates the ability of multimodal large language models to infer dyadic familiarity from short video clips of ice-breaker conversations. Across text, audio, and video modalities, the best model matches human accuracy, but exhibits a bias toward predicting 'stranger', unlike the balanced human responses. The benchmark reveals that while both humans and models benefit from richer modalities, only humans improve when visual behavior is added to speech.
The study shows that current multimodal LLMs can achieve human-level accuracy in social perception tasks but rely on different decision strategies, as evidenced by a skewed prior toward the 'stranger' class. This suggests that models may not be capturing the same social cues as humans, particularly from visual behavior, indicating a gap in true social understanding.
The release of FriendBench provides a standardized tool for evaluating social intelligence in AI systems, which is critical for applications in human-robot interaction, virtual assistants, and social media analysis. The finding that models and humans perform similarly but with different biases highlights the need for careful calibration in deployment.
For companies developing AI for social applications, FriendBench offers a way to measure and improve social perception capabilities, potentially leading to more natural and effective user interactions in products like virtual assistants, customer service bots, and social robots.
Future work may focus on improving models' ability to leverage visual social cues and reducing class imbalance biases. The benchmark could be expanded to more diverse social contexts and longer interactions, driving progress toward more socially aware AI.