arXiv:2510.21311cs.CV2025-10NeurIPS被引 3

用强化学习提升大模型对小物体的细粒度理解与分割能力

FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning

  • 分两阶段:先全局推理定位,再局部精细优化
  • 在4k分辨率图像上小物体分割精度超越现有方法
  • 适合需要精准识别微小目标的视觉任务研究者

多模态大语言模型在视觉-语言任务中表现卓越,但受限于输入分辨率,在高分辨率图像中难以精确理解与定位细节,尤其面对复杂背景下极小物体时。为此,我们提出 extsc{FineRS},一种基于强化学习的两阶段框架,联合实现极小物体的细粒度推理与分割。该框架采用从粗到精的流程:全局语义探索(GSE)生成文本响应和粗略目标区域;局部感知精炼(LPR)在此基础上输出精确边界框与分割掩码。通过引入定位感知的回溯奖励机制,将LPR输出用于优化GSE,增强粗略区域探索的鲁棒性。此外,我们构建了 extsc{FineRS}-4k 数据集,用于评估模型在复杂高分辨率场景中对细微小目标进行属性级推理与像素级分割的能力。在 extsc{FineRS}-4k 及公开数据集上的实验表明,该方法在指令引导分割与视觉推理任务中均持续优于当前最优的多模态大模型方法。

原文摘要 · Abstract (English)

Multi-modal Large Language Models (MLLMs) have shown remarkable capabilities across a wide range of vision-language tasks. However, due to the restricted input resolutions, MLLMs face significant challenges in precisely understanding and localizing visual details in high-resolution images -- particularly when dealing with extra-small objects embedded in cluttered contexts. To address this issue, we propose \textsc{FineRS}, a two-stage MLLM-based reinforcement learning framework for jointly reasoning and segmenting extremely small objects within high-resolution scenes. \textsc{FineRS} adopts a coarse-to-fine pipeline comprising Global Semantic Exploration (GSE) and Localized Perceptual Refinement (LPR). Specifically, GSE performs instruction-guided reasoning to generate a textural response and a coarse target region, while LPR refines this region to produce an accurate bounding box and segmentation mask. To couple the two stages, we introduce a locate-informed retrospective reward, where LPR's outputs are used to optimize GSE for more robust coarse region exploration. % Additionally, we present \textsc{FineRS}-4k, a new dataset for evaluating MLLMs on attribute-level reasoning and pixel-level segmentation on subtle, small-scale targets in complex high-resolution scenes. Experimental results on \textsc{FineRS}-4k and public datasets demonstrate that our method consistently outperforms state-of-the-art MLLM-based approaches on both instruction-guided segmentation and visual reasoning tasks.

小物体分割多模态大模型强化学习细粒度推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。