让智能体‘想象’多视角,提升视觉语言模型对人物交互的理解能力
What if Agents Could Imagine? Reinforcing Open-Vocabulary HOI Comprehension through Generation
- 用生成式框架让智能体模拟不同视角,弥补图像视角不足
- 在两个数据集上达到顶尖效果,训练数据仅需其他方法的36.7%
- 适合关注多模态推理与少样本学习的研究者
多模态大模型虽在跨模态推理中表现优异,但在开放词汇人-物交互(OV-HOI)任务中仍受限于跨模态幻觉和图像视角单一。为此,我们提出ImagineAgent,一种融合认知映射、工具增强强化学习与生成世界建模的智能体框架。首先构建名为hicodet-6K的思维链数据集,用于监督微调,通过结构化感知实体对实现全面预测。其次,设计多模态工具库,集成在线检索、图像裁剪与生成建模,使智能体能在推理中动态调用工具以缓解视觉-语义模糊与幻觉。此外,引入生成模型重构替代视角,使智能体能“想象”缺失视角下的场景。最后,采用复合奖励机制联合优化预测准确率与工具使用效率。在SWIG-HOI与HICO-DET数据集上的评估显示,本方法达到当前最优性能,且训练数据量仅为现有方法的36.7%,验证了其鲁棒性、有效性与高效性。
原文摘要 · Abstract (English)
Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Open-Vocabulary Human-Object Interaction (OV-HOI) are limited by cross-modal hallucinations and limited viewpoints of images. To address this, we propose ImagineAgent, an agentic framework that integrates cognitive mapping, tool-augmented reinforcement learning (RL), and generative world modeling for robust OV-HOI understanding. Specifically, we first propose an innovative CoT dataset named hicodet-6K for supervised fine-tuning (SFT), which effectively bridges the perception-to-cognition gap by structuring perceived entities into interaction pairs for comprehensive predictions. Subsequently, we develop a multimodal tool library integrating online retrieval, image cropping, and generative modeling, enabling the agent to dynamically augment reasoning with domain-specific tools to resolve visual-semantic ambiguities and hallucinations during inference. Moreover, we incorporate a generative model to reconstruct alternative viewpoints, enabling the agent to 'imagine' under limited viewpoints. Finally, we propose a composite reward mechanism to jointly optimize prediction accuracy and tool efficiency. Evaluations on both SWIG-HOI and HICO-DET datasets demonstrate that our method achieves state-of-the-art performance while requiring merely 36.7% of the training data compared to existing methods, validating our robustness, empirical effectiveness and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。