arXiv:2411.18038cs.CVcs.AI2024-11ECCV被引 4

用视觉语言模型提升人物交互分析的可解释性与准确率

VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis

  • 将视觉语言模型的语义匹配能力作为人物交互检测的目标函数
  • 在多个基准上达到当前最优性能,显著提升交互识别精度
  • 适合关注可解释性与多模态理解的研究者与应用开发者

大型视觉语言模型(VLM)在融合视觉与语言两大模态方面取得显著进展,具备对两种模态的全面理解能力,可完成多样化任务。本文提出一种新方法——VLM-HOI,首次将VLM的语言理解能力显式用于人-物交互(HOI)检测任务中。通过将预测的HOI三元组以自然语言形式表达,利用VLM进行图像-文本匹配评分,构建对比优化目标。相比CLIP等模型,该方法因具备定位能力和对象中心特性更适用于此类任务。实验表明,该方法在多个基准数据集上均实现最先进的检测精度,验证了其有效性。我们认为,将VLM引入HOI检测是迈向更高级、可解释的人物交互分析的重要一步。

原文摘要 · Abstract (English)

The Large Vision Language Model (VLM) has recently addressed remarkable progress in bridging two fundamental modalities. VLM, trained by a sufficiently large dataset, exhibits a comprehensive understanding of both visual and linguistic to perform diverse tasks. To distill this knowledge accurately, in this paper, we introduce a novel approach that explicitly utilizes VLM as an objective function form for the Human-Object Interaction (HOI) detection task (\textbf{VLM-HOI}). Specifically, we propose a method that quantifies the similarity of the predicted HOI triplet using the Image-Text matching technique. We represent HOI triplets linguistically to fully utilize the language comprehension of VLMs, which are more suitable than CLIP models due to their localization and object-centric nature. This matching score is used as an objective for contrastive optimization. To our knowledge, this is the first utilization of VLM language abilities for HOI detection. Experiments demonstrate the effectiveness of our method, achieving state-of-the-art HOI detection accuracy on benchmarks. We believe integrating VLMs into HOI detection represents important progress towards more advanced and interpretable analysis of human-object interactions.

人物交互视觉语言模型可解释性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。