arXiv:2604.10397cs.CVcs.AI2026-04被引 1

提出统一检测与预测的视频人物物体交互理解新框架

Rethinking Video Human-Object Interaction: Set Prediction over Time for Unified Detection and Anticipation

  • 以配对为中心建模,联合优化当前交互检测与未来演化预测
  • 在长时序预测上表现更优,比传统方法提升显著
  • 适用于需要理解动态交互场景的智能系统开发

基于视频的人物-物体交互(HOI)理解需同时检测当前交互并预测其未来演变。现有方法通常将预测作为下游任务,依赖外部构建的物体对,限制了检测与预测间的联合推理。此外,现有基准中稀疏的关键帧标注会导致未来标签与实际动态时间错位,降低预测评估可靠性。为此,我们提出 DETAnt-HOI,一个基于 VidHOI 与 Action Genome 的时序修正基准,支持更真实的多时程评估;并设计 HOI-DA 框架,通过将未来交互建模为当前配对状态的残差变化,实现主体-对象定位、当前交互检测与未来预测的联合学习。实验表明,在检测与预测任务上均取得持续提升,尤其在长时预测阶段增益更大。结果表明,将预测与检测联合学习可作为配对级视频表征学习的结构约束,显著提升性能。基准与代码将公开。

原文摘要 · Abstract (English)

Video-based human-object interaction (HOI) understanding requires both detecting ongoing interactions and anticipating their future evolution. However, existing methods usually treat anticipation as a downstream forecasting task built on externally constructed human-object pairs, limiting joint reasoning between detection and prediction. In addition, sparse keyframe annotations in current benchmarks can temporally misalign nominal future labels from actual future dynamics, reducing the reliability of anticipation evaluation. To address these issues, we introduce DETAnt-HOI, a temporally corrected benchmark derived from VidHOI and Action Genome for more faithful multi-horizon evaluation, and HOI-DA, a pair-centric framework that jointly performs subject-object localization, present HOI detection, and future anticipation by modeling future interactions as residual transitions from current pair states. Experiments show consistent improvements in both detection and anticipation, with larger gains at longer horizons. Our results highlight that anticipation is most effective when learned jointly with detection as a structural constraint on pair-level video representation learning. Benchmark and code will be publicly available.

视频理解交互预测联合建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。