arXiv:2508.18753cs.CV2025-08中稿 · CVPR被引 4

新基准让视觉语言模型与专用模型公平比拼人-物交互检测能力

CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods

  • 设计多选题形式的统一评测集,避免标签不匹配导致误判
  • 大模型零样本表现优异,但多人并发动作识别仍弱
  • 适合对比通用模型与专用模型在复杂场景下的优劣

人-物交互(HOI)检测长期由专用模型主导,部分使用CLIP等早期视觉语言模型(VLM)。随着大型生成式VLM兴起,关键问题是:独立VLM能否在HOI检测上媲美专用方法?现有基准如HICO-DET要求精确标签匹配,未标注即判错,这对自由输出的VLM不公平,导致跨范式比较不可靠。为此,我们提出CrossHOI-Bench,一个包含显式正例和精心筛选负例的多选题式评测集,支持对VLM和专用模型的统一可靠评估。重点考察多人场景和细粒度交互区分等挑战性任务,以揭示两类模型的真实差异。实验表明,大VLM实现竞争力甚至更优的零样本性能,但在多个并发动作识别及目标人物归属判断上表现不佳;而专用方法在一般推理上较弱,但在多动作识别和角色定位方面更可靠。这一发现揭示了两类模型的互补优缺点,而旧基准因错误惩罚机制未能呈现此真相。

原文摘要 · Abstract (English)

HOI detection has long been dominated by task-specific models, sometimes with early vision-language backbones such as CLIP. With the rise of large generative VLMs, a key question is whether standalone VLMs can perform HOI detection competitively against specialized HOI methods. Existing benchmarks such as HICO-DET require exact label matching under incomplete annotations, so any unmatched prediction is marked wrong. This unfairly penalizes valid outputs, especially from less constrained VLMs, and makes cross-paradigm comparison unreliable. To address this limitation, we introduce CrossHOI-Bench, a multiple-choice HOI benchmark with explicit positives and curated negatives, enabling unified and reliable evaluation of both VLMs and HOI-specific models. We further focus on challenging scenarios, such as multi-person scenes and fine-grained interaction distinctions, which are crucial for revealing real differences between the two paradigms. Experiments show that large VLMs achieve competitive, sometimes superior, zero-shot performance, yet they struggle with multiple concurrent actions and with correctly assigning interactions to the target person. Conversely, HOI-specific methods remain weaker in general HOI reasoning but demonstrate stronger multi-action recognition and more reliable identification of which person performs which action. These findings expose complementary strengths and weaknesses of VLMs and HOI-specific methods, which existing benchmarks fail to reveal due to incorrect penalization.

HOI检测视觉语言模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。