arXiv:2512.17640cs.CV2025-12

让模型生成交互动作,突破传统分类限制。

Generative Human-Object Interaction Detection via Differentiable Cognitive Steering of Multi-modal LLMs

  • 将HOI检测改为生成式任务,用大模型理解开放词汇。
  • 设计轻量级引导模块,把视觉信息注入冻结的多模态大模型。
  • 兼顾分类精度与零样本泛化能力,适合复杂场景应用。

人体-物体交互(HOI)检测旨在定位人与物体对及其交互关系。现有方法在封闭世界假设下将任务视为小规模预定义动词集上的分类问题,难以泛化到真实场景中长尾或模糊的未见交互。尽管近期多模态大语言模型(MLLMs)具备丰富的世界知识以支持开放词汇理解,但因其微调计算成本过高,仍与现有HOI检测器脱节。为此,我们提出 extsc{GRASP-HO}——一种生成式推理与可调节感知框架,将HOI检测从封闭集分类转变为开放词汇生成问题。为连接视觉与认知,我们首先提取混合交互表示,再设计轻量级可学习认知引导通道(CSC)模块,将细粒度视觉证据注入冻结的MLLM以实现有效推理。为缓解基于分类的HOI数据集与生成式模型间的监督不匹配问题,我们引入混合引导策略,结合语言建模损失与辅助分类损失,实现在不牺牲生成灵活性的前提下具备判别性定位能力。实验表明,该方法在封闭集上达到领先性能,并展现出强大的零样本泛化能力,实现了判别感知与生成推理的统一范式,适用于开放世界下的HOI检测。

原文摘要 · Abstract (English)

Human-object interaction (HOI) detection aims to localize human-object pairs and the interactions between them. Existing methods operate under a closed-world assumption, treating the task as a classification problem over a small, predefined verb set, which struggles to generalize to the long-tail of unseen or ambiguous interactions in the wild. While recent multi-modal large language models (MLLMs) possess the rich world knowledge required for open-vocabulary understanding, they remain decoupled from existing HOI detectors since fine-tuning them is computationally prohibitive. To address these constraints, we propose \GRASP-HO}, a novel Generative Reasoning And Steerable Perception framework that reformulates HOI detection from the closed-set classification task to the open-vocabulary generation problem. To bridge the vision and cognitive, we first extract hybrid interaction representations, then design a lightweight learnable cognitive steering conduit (CSC) module to inject the fine-grained visual evidence into a frozen MLLM for effective reasoning. To address the supervision mismatch between classification-based HOI datasets and open-vocabulary generative models, we introduce a hybrid guidance strategy that coupling the language modeling loss and auxiliary classification loss, enabling discriminative grounding without sacrificing generative flexibility. Experiments demonstrate state-of-the-art closed-set performance and strong zero-shot generalization, achieving a unified paradigm that seamlessly bridges discriminative perception and generative reasoning for open-world HOI detection.

HOI检测生成模型多模态开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。