arXiv:2605.16460cs.CV2026-05

通过优化推理过程提升视觉计数准确率

REC-RL: Referring expression counting via Gaussian and range-based reward optimization

论文配图:REC-RL: Referring expression counting via Gaussian and range-based reward optimization
图 1 · 摘自论文原文
  • 引入思考-范围-回答范式,优化中间推理步骤
  • 在多个基准上超越主流方法,提升推理质量
  • 适合关注视觉语言模型推理机制的研究者

指称表达计数(REC)是一项以意图为导向的任务,要求具备上下文感知的视觉推理能力。尽管现有视觉语言模型已融合语言理解视觉内容,但多数REC方法依赖基于规则的强化学习,奖励函数仅关注最终准确性,忽视中间推理质量。本文提出REC-RL框架,采用思考-范围-回答范式,显式优化视觉推理过程。该方法结合组相对策略优化(Group Relative Policy Optimization),设计两种轻量级奖励:基于区间监督与高斯精度引导的准确率奖励,以及强制结构化输出的格式奖励。通过将中间注意力预测建模为内部决策,避免额外标注,更贴近人类认知。大量实验表明,REC-RL在多个基准上持续优于强基线,展现出良好的泛化能力。

原文摘要 · Abstract (English)

Referring expression counting (REC) is an intention-driven task that requires context-aware visual reasoning. While recent vision-language models incorporate language for visual understanding, most existing REC methods rely on rulebased reinforcement learning with rewards focused primarily on final accuracy, overlooking the quality of intermediate reasoning. We propose REC-RL, a reinforcement learning framework that introduces a think-range-answer paradigm to explicitly optimize the visual reasoning process. RECRL employs Group Relative Policy Optimization and two lightweight rewards: an accuracy reward that combines range-based interval supervision with Gaussian-based precision guidance, and a format reward that enforces structured outputs. By modeling intermediate focus prediction as internal decision-making, REC-RL avoids additional annotations and better aligns with human perception. Extensive experiments demonstrate consistent improvements over strong baselines and robust generalization across benchmarks.

视觉推理强化学习计数任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。