arXiv:2505.19503cs.CV2025-05CVPR被引 17

提升CLIP对人物交互的细粒度感知能力,实现更精准零样本检测。

Locality-Aware Zero-Shot Human-Object Interaction Detection

  • 通过邻域特征聚合增强物体空间结构感知
  • 捕捉人与物体间的交互模式,提升识别准确率
  • 适合需要零样本泛化能力的视觉理解任务

零样本人类-物体交互(HOI)检测近年依赖大规模视觉-语言模型(如CLIP)在未见类别上的泛化能力,表现优异。然而现有方法难以适配CLIP对人-物对的表示,因其忽略区分交互所需的细粒度信息。为此,本文提出LAIN框架,通过引入局部性意识和交互意识来增强CLIP表示。局部性意识通过聚合邻近图像块的信息与空间先验,捕捉物体的细粒度细节与空间结构;交互意识则通过建模人与物体之间的交互模式,判断其是否互动及互动方式。实验证明,LAIN在多个基准上优于先前方法,在不同零样本设置下均表现出色,验证了局部性与交互意识对高效零样本HOI检测的重要性。

原文摘要 · Abstract (English)

Recent methods for zero-shot Human-Object Interaction (HOI) detection typically leverage the generalization ability of large Vision-Language Model (VLM), i.e., CLIP, on unseen categories, showing impressive results on various zero-shot settings. However, existing methods struggle to adapt CLIP representations for human-object pairs, as CLIP tends to overlook fine-grained information necessary for distinguishing interactions. To address this issue, we devise, LAIN, a novel zero-shot HOI detection framework enhancing the locality and interaction awareness of CLIP representations. The locality awareness, which involves capturing fine-grained details and the spatial structure of individual objects, is achieved by aggregating the information and spatial priors of adjacent neighborhood patches. The interaction awareness, which involves identifying whether and how a human is interacting with an object, is achieved by capturing the interaction pattern between the human and the object. By infusing locality and interaction awareness into CLIP representation, LAIN captures detailed information about the human-object pairs. Our extensive experiments on existing benchmarks show that LAIN outperforms previous methods on various zero-shot settings, demonstrating the importance of locality and interaction awareness for effective zero-shot HOI detection.

零样本检测视觉语言模型人物交互细粒度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。