arXiv:2607.13881cs.CVcs.AI2026-07被引 1

不训练模型,用多模态大模型实现真实场景下的人物物体交互检测。

Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild

论文配图:Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
图 1 · 摘自论文原文
  • 利用多模态大模型的推理能力,分步进行语义和空间定位分析。
  • 在真实复杂场景中实现全面且组合式的交互发现,准确率超越现有方法。
  • 适合需要快速部署、无需标注数据的开放世界交互检测任务。

人-物交互检测(HOID)传统上作为预定义交互类别的监督检测问题处理。尽管这类方法在封闭集基准上表现良好,但其将交互理解与特定数据集的监督绑定,限制了在开放世界和组合场景下的泛化能力。近期方法尝试通过提示策略利用多模态大语言模型(MLLM)迁移交互知识,但主要聚焦于提取判别性特征,忽视了其固有的多模态推理能力,导致在模糊和开放世界场景中缺乏有效上下文推理。本文提出AgentHOI,一种无需训练的代理式框架,将基础模型的通用多模态推理能力迁移到真实场景中的HOI检测。AgentHOI不学习交互分类器,而是模块化协同视觉基础模块,进行开放式语义推理与空间定位。针对复杂场景中交互发现不全与定位模糊的问题,引入两项关键机制:(1) 上下文感知多轮推理,逐步优化交互假设以确保全面且组合的发现;(2) 多维度交互定位,通过融合语义、空间与外观线索生成实例级描述,提升定位精度。大量实验表明,尽管无需任何HOI训练数据,AgentHOI在真实场景中性能优于最先进的监督与弱监督方法。

原文摘要 · Abstract (English)

Human-object interaction detection (HOID) has traditionally been formulated as a supervised detection problem over predefined interaction categories. While such paradigms achieve strong performance on closed-set benchmarks, they fundamentally entangle interaction understanding with dataset-specific supervision, limiting their ability to generalize to open-world and compositional scenarios. Recent HOI detectors attempt to leverage MLLMs through prompting strategies to transfer interaction-specific knowledge. However, such prompt-based approaches primarily focus on extracting discriminative representations from pretrained models, while underexploring their inherent multimodal reasoning capabilities. As a result, they struggle to provide informative contextual reasoning for ambiguous and open-world interaction scenarios. In this work, we present AgentHOI, a training-free, agentic framework that transfers the generalist multimodal reasoning capabilities of foundation models to HOI detection in the wild. Instead of learning interaction classifiers, AgentHOI modularly orchestrates complementary vision foundation modules to perform open-ended semantic reasoning and spatial grounding in a coordinated manner. To address the challenges of incomplete interaction discovery and ambiguous localization in complex scenes, we introduce two key mechanisms: (1) Context-aware Multi-round Reasoning, which progressively refines interaction hypotheses to ensure exhaustive and compositional HOI discovery, and (2) Multifaceted Interaction Localization, which enhances grounding precision by generating instance-specific descriptions that integrate semantic, spatial, and appearance cues. Extensive experiments demonstrate that AgentHOI achieves superior performance over state-of-the-art supervised and weakly supervised methods in real-world settings, despite requiring no HOID data for training.

交互检测多模态大模型零样本开放世界

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。