arXiv:2410.20155cs.CV2024-10NeurIPS被引 29

用扩散模型提升人物交互检测,零样本表现更优

Human-Object Interaction Detection Collaborated with Large Relation-driven Diffusion Models

  • 基于扩散模型学习关系嵌入,生成交互图像以提取线索
  • 在三个数据集上均达当前最优,零样本场景下效果突出
  • 适合研究视觉-语言对齐与零样本推理的学者

主流的人物交互(HOI)检测方法依赖大规模视觉-语言模型识别涉及人类和物体的事件。尽管前景可观,但通过图文对比学习训练的模型常忽略中低层视觉特征,且在组合推理方面表现不佳。为此,我们提出DIFFUSIONHOI,一种利用文本到图像扩散模型的新式HOI检测器。不同于传统模型,扩散模型作为生成模型擅长捕捉中低层视觉概念,并具备强组合性,可处理文本输入中的新概念。鉴于扩散模型通常聚焦实例物体,我们首先设计一种基于反演的策略,在嵌入空间中学习人类与物体间的关系模式表达。这些学习到的关系嵌入作为文本提示,引导扩散模型生成体现特定交互的图像,并从中提取与HOI相关的线索,无需大量微调。得益于上述机制,DIFFUSIONHOI在三个数据集上均实现当前最优性能,涵盖常规与零样本设置。

原文摘要 · Abstract (English)

Prevalent human-object interaction (HOI) detection approaches typically leverage large-scale visual-linguistic models to help recognize events involving humans and objects. Though promising, models trained via contrastive learning on text-image pairs often neglect mid/low-level visual cues and struggle at compositional reasoning. In response, we introduce DIFFUSIONHOI, a new HOI detector shedding light on text-to-image diffusion models. Unlike the aforementioned models, diffusion models excel in discerning mid/low-level visual concepts as generative models, and possess strong compositionality to handle novel concepts expressed in text inputs. Considering diffusion models usually emphasize instance objects, we first devise an inversion-based strategy to learn the expression of relation patterns between humans and objects in embedding space. These learned relation embeddings then serve as textual prompts, to steer diffusion models generate images that depict specific interactions, and extract HOI-relevant cues from images without heavy fine-tuning. Benefited from above, DIFFUSIONHOI achieves SOTA performance on three datasets under both regular and zero-shot setups.

人机交互扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。