让大模型统一理解视觉显著性,提升关键物体识别能力。
Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
- 用结构化标签统一处理三类显著性任务,实现一模多用。
- 在三个任务上均超越或持平顶尖模型,显著性推理能力更强。
- 提出新型强化学习算法,训练更高效、信号更稳定,适合研究者参考。
尽管多模态大语言模型在高层视觉-语言推理上表现优异,但缺乏对视觉显著性的内在感知,难以识别关键视觉元素。为弥补这一差距,我们提出Saliency-R1,首个统一的MLLM框架,联合解决三种典型且异构的显著性任务:显著性物体检测(SOD)、显著性实例分割(SIS)和共显著性物体检测(CoSOD),增强模型的显著性推理能力。我们引入带有结构化标签(<rg>、<ins>)的文本接口,编码区域与实例级指代表达,使单一指代分割器可生成适配不同任务的掩码。为高效训练MLLM,我们提出置信度引导策略优化(CGPO),一种新型单样本强化学习算法。CGPO通过用基于奖励-置信度差异的样本级信号替代组归一化优势,减少计算浪费,缓解信号稀释,降低训练开销。我们的模型在所有三项任务中均超越或持平主流开源/闭源MLLM及专用最先进方法,验证了该框架在显著性推理上的有效性。
原文摘要 · Abstract (English)
Although multimodal large language models (MLLMs) excel in high-level vision-language reasoning, they lack inherent awareness of visual saliency, making it difficult to identify key visual elements. To bridge this gap, we propose Saliency-R1, the first unified MLLM framework that jointly tackles three representative and heterogeneous saliency tasks: Salient Object Detection (SOD), Salient Instance Segmentation (SIS), and Co-salient Object Detection (CoSOD), enhancing the model's capacity for saliency reasoning. We introduce a textual interface with structured tags (<rg>, <ins>) to encode region- and instance-level referring expressions, enabling a single referring segmenter to produce task-appropriate masks. To train the MLLM efficiently, we propose Confidence-Guided Policy Optimization (CGPO), a novel single-sample reinforcement learning algorithm. CGPO improves on GRPO by replacing group-normalized advantages with a per-sample signal based on reward-confidence discrepancy, thereby reducing computational waste, mitigating signal dilution, and lowering training overhead. Our model exceeds or matches the performance of robust open/closed-source MLLMs and specialized state-of-the-art methods across all three tasks, demonstrating the efficacy of our framework in saliency reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。