arXiv:2507.12628cs.CV2025-07被引 1

提出顶向下框架Funnel-HOI,提升零样本人-物交互检测性能。

Funnel-HOI: Top-Down Perception for Zero-Shot HOI Detection

  • 先探测物体再关联动作,利用多模态共注意力挖掘交互线索。
  • 在HICO-DET和V-COCO上,未见类别最高提升12.4%,稀有类别提升8.4%。
  • 适合关注零样本视觉理解与交互检测的开发者和研究者。

人-物交互检测(HOID)旨在定位图像中的人物交互对并识别其交互行为。由于可能存在指数级的物-动组合,标注数据有限,导致长尾分布问题。近年来,零样本学习成为解决方案,基于端到端变换器的物体检测器已成功应用于HOID。然而,现有方法主要聚焦于改进解码器以学习交互的纠缠或解纠缠表示。本文主张应在编码器阶段就预判特定于HOI的线索,以获得更强的场景理解。为此,我们提出了名为Funnel-HOI的顶向下框架,受人类在场景理解中先把握明确概念再关联抽象概念的启发。首先探测图像中的物体(明确概念),再探测与之相关的动作(抽象概念)。一种新颖的非对称共注意力机制利用多模态信息(包含零样本能力),在编码器层面生成更强的交互表示。此外,设计了一种新损失函数,考虑物-动相关性,比现有损失更有效地调节误分类惩罚,指导交互分类器。在HICO-DET和V-COCO数据集上的大量实验表明,在全监督及六种零样本设置下均达到当前最优性能,对未见和稀有交互类别分别提升12.4%和8.4%。

原文摘要 · Abstract (English)

Human-object interaction detection (HOID) refers to localizing interactive human-object pairs in images and identifying the interactions. Since there could be an exponential number of object-action combinations, labeled data is limited - leading to a long-tail distribution problem. Recently, zero-shot learning emerged as a solution, with end-to-end transformer-based object detectors adapted for HOID becoming successful frameworks. However, their primary focus is designing improved decoders for learning entangled or disentangled interpretations of interactions. We advocate that HOI-specific cues must be anticipated at the encoder stage itself to obtain a stronger scene interpretation. Consequently, we build a top-down framework named Funnel-HOI inspired by the human tendency to grasp well-defined concepts first and then associate them with abstract concepts during scene understanding. We first probe an image for the presence of objects (well-defined concepts) and then probe for actions (abstract concepts) associated with them. A novel asymmetric co-attention mechanism mines these cues utilizing multimodal information (incorporating zero-shot capabilities) and yields stronger interaction representations at the encoder level. Furthermore, a novel loss is devised that considers objectaction relatedness and regulates misclassification penalty better than existing loss functions for guiding the interaction classifier. Extensive experiments on the HICO-DET and V-COCO datasets across fully-supervised and six zero-shot settings reveal our state-of-the-art performance, with up to 12.4% and 8.4% gains for unseen and rare HOI categories, respectively.

零样本交互检测顶向下多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。