用强化学习让大模型更懂人物交互的时空逻辑。
HOID-R1: Reinforcement Learning for Open-World Human-Object Interaction Detection Reasoning with Multimodal Large Language Model
- 结合思维链与强化学习,让模型边推理边优化
- 在多个数据集上达到当前最佳性能,泛化能力更强
- 适合做开放世界视觉理解的科研与工程人员
理解与识别人物交互(HOI)是AR/VR和机器人领域的重要应用。现有开放词汇HOI检测方法依赖大语言模型生成丰富文本提示,却忽视了其内在的3D空间理解能力。为此,我们提出HOID-R1,首个将思维链(CoT)引导的监督微调(SFT)与组相对策略优化(GRPO)结合于强化学习框架中的HOI检测方法。首先通过SFT赋予模型必要的推理能力,强制输出包含思考过程;随后引入GRPO,利用多奖励信号优化策略,增强跨模态对齐。为缓解思维链推理中的幻觉问题,设计了“多模态大模型作为裁判”机制,监督推理输出,进一步提升泛化性。大量实验表明,HOID-R1在多个HOI检测基准上达到领先性能,并在开放世界新场景下表现出更强的泛化能力。
原文摘要 · Abstract (English)
Understanding and recognizing human-object interaction (HOI) is a pivotal application in AR/VR and robotics. Recent open-vocabulary HOI detection approaches depend exclusively on large language models for richer textual prompts, neglecting their inherent 3D spatial understanding capabilities. To address this shortcoming, we introduce HOID-R1, the first HOI detection framework that integrates chain-of-thought (CoT) guided supervised fine-tuning (SFT) with group relative policy optimization (GRPO) within a reinforcement learning (RL) paradigm. Specifically, we initially apply SFT to imbue the model with essential reasoning capabilities, forcing the model to articulate its thought process in the output. Subsequently, we integrate GRPO to leverage multi-reward signals for policy optimization, thereby enhancing alignment across diverse modalities. To mitigate hallucinations in the CoT reasoning, we introduce an "MLLM-as-a-judge" mechanism that supervises the CoT outputs, further improving generalization. Extensive experiments show that HOID-R1 achieves state-of-the-art performance on HOI detection benchmarks and outperforms existing methods in open-world generalization to novel scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。