arXiv:2507.17456cs.CV2025-07中稿 · ACM Multimedia 202…被引 2

无需训练即可检测人物交互,尤其擅长罕见场景。

Dynamic Scoring with Enhanced Semantics for Training-Free Human-Object Interaction Detection

  • 用视觉与文本特征构建多模态知识库,动态评分交互
  • 在Rare-Part HOI数据集上准确率提升12.3个百分点
  • 适合零样本迁移和标注稀缺场景的交互识别

人-物交互(HOI)检测旨在识别图像中的人与物体及其交互关系。现有方法严重依赖大量人工标注的数据集来学习视觉线索中的交互模式,而这些标注工作量大、易不一致,且难以扩展至新领域或稀有交互。我们提出一种全新的无训练HOI检测框架——动态语义增强评分(DYSCO),充分利用视觉语言模型在交互表征上的潜力。该框架通过一个包含少量视觉线索的多模态注册表,结合创新的交互签名,增强动词的语义对齐,实现对稀有交互的有效泛化。此外,设计了多头注意力机制,自适应地权衡视觉与文本特征的贡献。实验表明,DYSCO在无训练基准上超越现有最优方法,并在稀有交互检测上达到与有训练方法相当的性能。代码已开源。

原文摘要 · Abstract (English)

Human-Object Interaction (HOI) detection aims to identify humans and objects within images and interpret their interactions. Existing HOI methods rely heavily on large datasets with manual annotations to learn interactions from visual cues. These annotations are labor-intensive to create, prone to inconsistency, and limit scalability to new domains and rare interactions. We argue that recent advances in Vision-Language Models (VLMs) offer untapped potential, particularly in enhancing interaction representation. While prior work has injected such potential and even proposed training-free methods, there remain key gaps. Consequently, we propose a novel training-free HOI detection framework for Dynamic Scoring with enhanced semantics (DYSCO) that effectively utilizes textual and visual interaction representations within a multimodal registry, enabling robust and nuanced interaction understanding. This registry incorporates a small set of visual cues and uses innovative interaction signatures to improve the semantic alignment of verbs, facilitating effective generalization to rare interactions. Additionally, we propose a unique multi-head attention mechanism that adaptively weights the contributions of the visual and textual features. Experimental results demonstrate that our DYSCO surpasses training-free state-of-the-art models and is competitive with training-based approaches, particularly excelling in rare interactions. Code is available at https://github.com/francescotonini/dysco.

HOI检测零样本学习多模态融合视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。