arXiv:2604.14069cs.CV2026-04中稿 · the 20th IEEE Inte…

让模型在野外识别任意人物交互,无需预先定义动作列表。

Towards Unconstrained Human-Object Interaction

论文配图:Towards Unconstrained Human-Object Interaction
图 1 · 摘自论文原文
  • 用多模态大模型直接解析自由文本描述的交互关系。
  • 首次提出无约束人物交互任务,突破预设动作集限制。
  • 适合研究开放场景下视觉理解与语言模型融合的学者。

人-物交互(HOI)检测是计算机视觉中的长期难题,旨在预测人与物体之间的交互行为。当前的HOI模型依赖于训练和推理时预设的交互词汇表,限制了其在动态环境中的应用。随着多模态大语言模型(MLLMs)的发展,探索更灵活的交互识别范式成为可能。本文从MLLM视角重新审视HOI检测,提出在真实场景中进行无约束人-物交互(U-HOI)检测的新任务,该任务在训练和推理阶段均无需预先定义交互类别。我们评估了多种MLLMs在此设定下的表现,并引入一个包含测试时推理和语言到图转换的管道,从自由文本中提取结构化交互信息。实验表明,现有HOI检测器存在明显局限,而MLLMs在处理U-HOI任务中展现出显著价值。代码将公开于 https://github.com/francescotonini/anyhoi。

原文摘要 · Abstract (English)

Human-Object Interaction (HOI) detection is a longstanding computer vision problem concerned with predicting the interaction between humans and objects. Current HOI models rely on a vocabulary of interactions at training and inference time, limiting their applicability to static environments. With the advent of Multimodal Large Language Models (MLLMs), it has become feasible to explore more flexible paradigms for interaction recognition. In this work, we revisit HOI detection through the lens of MLLMs and apply them to in-the-wild HOI detection. We define the Unconstrained HOI (U-HOI) task, a novel HOI domain that removes the requirement for a predefined list of interactions at both training and inference. We evaluate a range of MLLMs on this setting and introduce a pipeline that includes test-time inference and language-to-graph conversion to extract structured interactions from free-form text. Our findings highlight the limitations of current HOI detectors and the value of MLLMs for U-HOI. Code will be available at https://github.com/francescotonini/anyhoi

人物交互多模态开放域大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。