arXiv:2607.29602cs.CLcs.AI2026-07

评测人类与大模型从20秒对话推断关系亲疏的能力,发现顶尖模型与人表现相当但倾向更保守。

FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models

  • 用20秒双人破冰对话片段,仅凭互动方式判断两人是否相识
  • 最强模型与人类在文本/音频/视频三模态上准确率相当,均超85%
  • 人类能更好利用视觉行为线索,模型则更依赖语言,存在不同先验偏好

社交情境的理解常依赖行为而非言语。我们提出 FriendBench,一个基于20秒双人破冰对话片段的基准,用于判断两人是已相识还是初次见面。所有配对回答相同提示,仅通过互动方式揭示答案。我们在文本、音频和视频三个模态下,对比了来自七家公司的26个模型与匹配的人类小组在96组平衡双人组合上的表现。最佳模型与人类群体在各模态下的准确率无统计差异,但实现路径不同:人类在两种答案间保持平衡,而最强模型倾向于“陌生人”——反映的是有效先验差异,而非辨别能力。更丰富的信息通道对两者帮助不均等,且仅人类能从语音之外的视觉行为中获益。我们公开了刺激数据、人类评分及模型预测结果。

原文摘要 · Abstract (English)

Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models lean toward "stranger"---a difference in effective prior, not discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.

多模态关系推理基准测试人类对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。