arXiv:2512.00885cs.CV2025-12被引 2

构建细粒度手物交互动态视频问答基准,提升模型对动作与变化的精准理解。

HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics

  • 设计六类问题覆盖操作与效果,实现时空层面的细粒度推理。
  • 包含11.1K个问答对与10.3K个分割掩码,支持部件级语义分析。
  • 揭示当前模型在空间关系与运动理解上的不足,适合视觉-语言模型研究者。

手物交互(HOI)本质上涉及动态过程,人类操作会在时空上对物体产生特定影响。然而,现有语义HOI基准要么聚焦于操作,要么仅粗粒度关注结果,缺乏对交互动态的细粒度时空推理能力。我们提出HanDyVQA,一个全面涵盖操作与效应两方面的细粒度视频问答基准。该数据集包含六类互补问题(动作、过程、物体、位置、状态变化、物体部件),共11.1K个多项选择题对,涵盖操纵风格、手/物运动及部件级状态变化识别。此外,还提供10.3K个物体与部件的分割掩码,支持视频对象分割中的部件级推理评估。我们在该基准上评测了多个主流视频基础模型,发现表现最佳的Gemini-2.5-Pro平均准确率仅为73%,远低于人类水平(97%)。进一步分析表明,模型在空间关系、运动理解和部件几何结构方面仍存在显著挑战。实验还显示,将显式的HOI相关提示融入视觉特征可有效提升性能,为未来模型设计提供了关键洞见。

原文摘要 · Abstract (English)

Hand-object interaction (HOI) inherently involves dynamics where human manipulations produce distinct spatio-temporal effects on objects. However, existing semantic HOI benchmarks focused either on manipulation or on the resulting effects at a coarse level, lacking fine-grained spatio-temporal reasoning to capture the underlying dynamics in HOI. We introduce HanDyVQA, a fine-grained video question-answering benchmark that comprehensively covers both the manipulation and effect aspects of HOI. HanDyVQA comprises six complementary question types (Action, Process, Objects, Location, State Change, and Object Parts), totalling 11.1K multiple-choice QA pairs. Collected QA pairs recognizing manipulation styles, hand/object motions, and part-level state changes. HanDyVQA also includes 10.3K segmentation masks for Objects and Object Parts questions, enabling the evaluation of object/part-level reasoning in video object segmentation. We evaluated recent video foundation models on our benchmark and found that even the best-performing model, Gemini-2.5-Pro, reached only 73% average accuracy, which is far from human performance (97%). Further analysis shows the remaining challenges in spatial relationship, motion, and part-level geometric understanding. We also found that integrating explicit HOI-related cues into visual features improves performance, offering insights for developing future models with a deeper understanding of HOI dynamics.

视频问答手物交互细粒度推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。