arXiv:2511.00396cs.CV2025-11

让大模型统一理解视觉显著性,提升关键物体识别能力。

Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning

  • 用结构化标签统一处理三类显著性任务,实现一模多用。
  • 在三个任务上均超越或持平顶尖模型,显著性推理能力更强。
  • 提出新型强化学习算法,训练更高效、信号更稳定,适合研究者参考。

尽管多模态大语言模型在高层视觉-语言推理上表现优异,但缺乏对视觉显著性的内在感知,难以识别关键视觉元素。为弥补这一差距,我们提出Saliency-R1,首个统一的MLLM框架,联合解决三种典型且异构的显著性任务:显著性物体检测(SOD)、显著性实例分割(SIS)和共显著性物体检测(CoSOD),增强模型的显著性推理能力。我们引入带有结构化标签(<rg>、<ins>)的文本接口,编码区域与实例级指代表达,使单一指代分割器可生成适配不同任务的掩码。为高效训练MLLM,我们提出置信度引导策略优化(CGPO),一种新型单样本强化学习算法。CGPO通过用基于奖励-置信度差异的样本级信号替代组归一化优势,减少计算浪费,缓解信号稀释,降低训练开销。我们的模型在所有三项任务中均超越或持平主流开源/闭源MLLM及专用最先进方法,验证了该框架在显著性推理上的有效性。

原文摘要 · Abstract (English)

Although multimodal large language models (MLLMs) excel in high-level vision-language reasoning, they lack inherent awareness of visual saliency, making it difficult to identify key visual elements. To bridge this gap, we propose Saliency-R1, the first unified MLLM framework that jointly tackles three representative and heterogeneous saliency tasks: Salient Object Detection (SOD), Salient Instance Segmentation (SIS), and Co-salient Object Detection (CoSOD), enhancing the model's capacity for saliency reasoning. We introduce a textual interface with structured tags (<rg>, <ins>) to encode region- and instance-level referring expressions, enabling a single referring segmenter to produce task-appropriate masks. To train the MLLM efficiently, we propose Confidence-Guided Policy Optimization (CGPO), a novel single-sample reinforcement learning algorithm. CGPO improves on GRPO by replacing group-normalized advantages with a per-sample signal based on reward-confidence discrepancy, thereby reducing computational waste, mitigating signal dilution, and lowering training overhead. Our model exceeds or matches the performance of robust open/closed-source MLLMs and specialized state-of-the-art methods across all three tasks, demonstrating the efficacy of our framework in saliency reasoning.

多模态显著性强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。