arXiv:2412.19542cs.CVcs.AI2024-12AAAI被引 6

提出新数据集GIO,解决视频中多样物体的交互定位难题

Interacted Object Grounding in Spatio-Temporal Human-Object Interactions

  • 设计4D问答框架,融合时空线索定位交互物体
  • 在29万标注框上实现显著优于基线的定位性能
  • 适合研究视频理解与开放世界物体识别的学者

时空人体-物体交互(ST-HOI)理解旨在从视频中检测人体与物体的交互,对活动理解至关重要。然而,现有全身体交互视频基准忽略了开放世界物体的多样性——通常只提供有限且预定义的物体类别。为此,我们提出一个全新的开放世界基准:交互物体定位(GIO),包含1,098个交互物体类别和290,000个交互物体边界框标注。基于此,我们提出了一个物体定位任务,要求视觉系统发现交互物体。尽管当前检测器与定位方法已取得显著进展,但在GIO中对多样化、稀有物体的定位仍表现不佳,深刻揭示了现有视觉系统的局限性并带来巨大挑战。因此,我们探索利用时空线索来解决物体定位问题,提出4D问答框架(4D-QA),从多样化视频中发现交互物体。大量实验表明,该方法在性能上显著优于现有基线。数据与代码将公开于 https://github.com/DirtyHarryLYL/HAKE-AVA。

原文摘要 · Abstract (English)

Spatio-temporal Human-Object Interaction (ST-HOI) understanding aims at detecting HOIs from videos, which is crucial for activity understanding. However, existing whole-body-object interaction video benchmarks overlook the truth that open-world objects are diverse, that is, they usually provide limited and predefined object classes. Therefore, we introduce a new open-world benchmark: Grounding Interacted Objects (GIO) including 1,098 interacted objects class and 290K interacted object boxes annotation. Accordingly, an object grounding task is proposed expecting vision systems to discover interacted objects. Even though today's detectors and grounding methods have succeeded greatly, they perform unsatisfactorily in localizing diverse and rare objects in GIO. This profoundly reveals the limitations of current vision systems and poses a great challenge. Thus, we explore leveraging spatio-temporal cues to address object grounding and propose a 4D question-answering framework (4D-QA) to discover interacted objects from diverse videos. Our method demonstrates significant superiority in extensive experiments compared to current baselines. Data and code will be publicly available at https://github.com/DirtyHarryLYL/HAKE-AVA.

视频理解物体定位开放世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。