首个评估视听对话系统全双工能力的基准,推动更自然的人机交互。
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents

- 构建首个视听到视听的全双工对话评测框架,涵盖237段真实视频通话片段。
- 发现现有模型存在视觉流忽略和描述坍缩问题,无法实现持续视听协同。
- 适合研究多模态对话、人机交互与具身智能的学者使用。
自然人类对话是全双工且视听结合的:人们在同时说话与倾听的同时,持续解读和生成非语言线索,如点头、微笑和手势。为实现高效人机交互,对话代理需建模全双工视听对话;然而现有全双工评测仅关注语音。本文提出VideoFDB,首个评估视听到视听(AV2AV)全双工对话代理的基准。VideoFDB包含:(i) 237段双人对话片段,覆盖11种真实视频通话中的非语言互动动态;(ii) 区分感知与生成行为的分类体系;(iii) 基于评分框架的LM-as-judge评估方法,可解释地评估对话质量与非语言互动动态。我们对开源与闭源视觉-语音代理进行评估,发现系统性失败模式:描述坍缩与视觉流忽视。当前系统仅将视觉用于显式视觉问答,而非自然对话所需的连续视听联合定位。进一步评估级联语音到形象系统,发现其架构从根本上无法生成全双工非语言线索。作为首个全双工AV2AV基准,VideoFDB为系统性评估奠定基础,有望加速下一代多模态对话代理的发展。
原文摘要 · Abstract (English)
Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smiles, and gestures. To support successful human-agent interaction, agents must model full-duplex audiovisual conversation; however, existing full-duplex benchmarks evaluate only speech. In this work, we present VideoFDB, the first benchmark to evaluate full-duplex audio-visual-to-audio-visual (AV2AV) conversational agents. VideoFDB contributes (i) 237 dyadic clips spanning 11 nonverbal conversational dynamics from real-world video calls, (ii) a taxonomy separating perception from generation behaviors, and (iii) a rubric-based LM-as-judge evaluation framework with interpretable axes for assessing conversational quality with respect to nonverbal conversational dynamics. Across open- and closed-source vision-speech agents, we find systematic failure modes: captioning collapse and visual-stream ignorance, and we show that current systems exploit vision for explicit visual question answering but not for the streaming joint audiovisual grounding required in natural conversation. We further evaluate cascaded speech-to-avatar systems and find that their architecture fundamentally precludes the production of full-duplex nonverbal cues. As the first benchmark for full-duplex AV2AV interaction, VideoFDB establishes a foundation for systematic evaluation and, we hope, will accelerate the advancement and development of next-generation multimodal conversational agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。