arXiv:2507.09313cs.CV2025-07被引 20

首个评估视频大模型主动交互能力的基准,推动更自然的人机对话。

ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models

  • 构建主动交互评测基准,模拟真实视频播放中的实时响应
  • 提出新指标PAUC,考虑响应时间动态性,更贴近用户偏好
  • 实验证明PAUC比传统文本指标更能反映真实用户体验

随着多模态对话系统研究的深入,主动交互能力日益受到重视。与传统的逐轮对话不同,用户期望系统能在视频播放过程中自主决定回应时机,实现更主动的多轮交互。为推动该领域发展,我们提出首个综合性评测基准ProactiveVideoQA,用于评估系统在主动交互场景下的表现。由于模型响应时间各不相同,我们进一步提出PAUC——首个考虑响应时间动态性的评价指标,使主动交互场景下的评估更精准。通过对多种基线系统的广泛评测及用户偏好研究,我们发现PAUC相较于仅关注文本内容的传统指标,与人类偏好更具一致性。结果表明,PAUC能更真实地衡量主动交互中的用户体验。

原文摘要 · Abstract (English)

With the growing research focus on multimodal dialogue systems, the capability for proactive interaction is gradually gaining recognition. As an alternative to conventional turn-by-turn dialogue, users increasingly expect multimodal systems to be more initiative, for example, by autonomously determining the timing of multi-turn responses in real time during video playback. To facilitate progress in this emerging area, we introduce ProactiveVideoQA, the first comprehensive benchmark to evaluate a system's ability to engage in proactive interaction. Since model responses are generated at varying timestamps, we further propose PAUC, the first metric that accounts for the temporal dynamics of model responses. This enables a more accurate evaluation of systems operating in proactive settings. Through extensive benchmarking of various baseline systems on ProactiveVideoQA and a user study of human preferences, we show that PAUC is in better agreement with human preferences than traditional evaluation metrics, which typically only consider the textual content of responses. These findings demonstrate that PAUC provides a more faithful assessment of user experience in proactive interaction scenarios. Project homepage: https://github.com/yellow-binary-tree/ProactiveVideoQA

视频问答主动交互评估基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。