arXiv:2608.01055cs.CVcs.AI2026-08

让每个预测框得到应得的奖励,提升多目标视觉定位精度

Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception

论文配图:Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception
图 1 · 摘自论文原文
  • 通过留一法计算每个框的边际贡献,精准分配奖励
  • 在多个基准上超越现有GRPO方法,计数与分割更准
  • 适合需要精细定位和多对象推理的任务研究者

多模态大语言模型日益被期待完成结构化感知任务,包括视觉识别、语言到物体的绑定、对象数量保持以及精确定位和分割输出。然而,现有的组内相对强化学习方法仅提供响应级监督,导致结构化多对象预测中的粒度不匹配:单一优势被广播至响应中所有标记,无法区分单个框的贡献。为解决这一问题,我们提出MCR-GRPO,一种从每条采样响应中直接推导框级信用的边际贡献分配框架。具体而言,边际贡献奖励(MCR)通过留一法比较,衡量当某个框被移除后匹配集合值的变化,从而估计其贡献。经响应内归一化后,提升集合值的记录获得正向信用,冗余或有害项则被抑制。为进一步提升边际归因的稳定性与信息量,我们引入连续匹配集合值评估器,集成排列不变匹配、计数感知归一化与分级定位机制。MCR-GRPO将归一化的框级边际优势映射至生成该框的标记跨度,保留GRPO的响应级比较能力,同时实现对结构化多对象定位的框级优化。在REC、DOD、分割和计数等基准上的实验表明,该方法在性能上达到当前最优水平。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) are increasingly expected to solve structured perception tasks that require visual recognition, language-to-object binding, object cardinality preservation, and precisely localized grounding and segmentation outputs. However, existing group-relative reinforcement learning methods provide only response-level supervision, creating a granularity mismatch for structured multi-object prediction: a single advantage is broadcast to all tokens in a response, without distinguishing individual box contributions. To address this mismatch, we propose MCR-GRPO, a marginal contribution assignment framework that derives box-level credit directly from each sampled response. Specifically, Marginal Contribution Reward (MCR) estimates each predicted box's contribution through a leave-one-out comparison, measuring how the matched set value changes when the box is removed from the response. After within-response normalization, records that improve the set value receive positive credit, while redundant or harmful ones are suppressed. To make marginal attribution stable and informative, we further introduce a Continuous Matched Set Value Evaluator that integrates permutation-invariant matching, count-aware normalization, and graded localization. MCR-GRPO maps normalized box-level marginal advantages to the token spans that generated each box, preserving GRPO's response-level comparison while enabling box-aware optimization of structured multi-object grounding. Experiments across REC, DOD, segmentation, and counting benchmarks show state-of-the-art performance over prior GRPO-based baselines.

多模态目标定位强化学习结构化感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。