arXiv:2504.14642cs.CV2025-04AAAI被引 8

首个统一关系理解框架,用思维链强化视觉语义关联。

Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension

  • 分步引导:先思维链微调,再多奖励优化
  • 在PSG和SWiG数据集上超越现有模型
  • 适合研究视觉关系推理与大模型对齐的学者

多模态大模型在物体定位和区域描述方面取得进展,但在视觉关系理解上仍受限,难以处理二元关系,更别提涉及多个语义角色的N元关系。根本原因在于缺乏对多实体间结构化语义依赖的建模,导致输出不可靠、幻觉频发,并过度依赖语言先验(如人拿着杯子就默认为‘喝牛奶’)。为此,我们提出Relation-R1,首个统一的关系理解框架,将思维链(CoT)引导的监督微调(SFT)与组相对策略优化(GRPO)结合于强化学习范式中。首先通过SFT建立结构化推理能力,强制生成带思考过程的输出;随后利用GRPO通过多奖励优化,优先考虑视觉-语义对齐而非语言偏差,提升泛化能力。进一步研究不同CoT策略,发现从具体到一般的渐进式引导可进一步提升泛化性,尤其在捕捉同义N元关系方面表现突出。在广泛使用的PSG和SWiG数据集上的实验表明,Relation-R1在二元和N元关系理解上均达到当前最优性能。

原文摘要 · Abstract (English)

Recent advances in multi-modal large language models (MLLMs) have significantly improved object-level grounding and region captioning. However, they remain limited in visual relation understanding, struggling even with binary relation detection, let alone \textit{N}-ary relations involving multiple semantic roles. The core reason is the lack of modeling for \textit{structural semantic dependencies} among multi-entities, leading to unreliable outputs, hallucinations, and over-reliance on language priors (\eg, defaulting to ``person drinks a milk'' if a person is merely holding it). To this end, we propose Relation-R1, the \textit{first unified} relation comprehension framework that explicitly integrates cognitive chain-of-thought (CoT)-guided supervised fine-tuning (SFT) and group relative policy optimization (GRPO) within a reinforcement learning (RL) paradigm. Specifically, we first establish foundational reasoning capabilities via SFT, enforcing structured outputs with thinking processes. Then, GRPO is utilized to refine these outputs via multi-rewards optimization, prioritizing visual-semantic grounding over language-induced biases, thereby improving generalization capability. Furthermore, we investigate the impact of various CoT strategies within this framework, demonstrating that a specific-to-general progressive approach in CoT guidance further improves generalization, especially in capturing synonymous \textit{N}-ary relations. Extensive experiments on widely-used PSG and SWiG datasets demonstrate that Relation-R1 achieves state-of-the-art performance in both binary and \textit{N}-ary relation understanding.

关系理解思维链强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。