arXiv:2605.07394cs.CVcs.AI2026-05

平衡多维度评价,提升大模型图像描述的准确性与流畅性

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning

论文配图:BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning
图 1 · 摘自论文原文
  • 设计多目标强化学习框架,同时优化准确度、覆盖性和语言质量
  • 在多个模型上实现最高13.6分的DCScore提升,显著改善生成质量
  • 适合关注图像描述质量与实用性的研究者和开发者

图像描述是计算机视觉中的基础任务。随着多模态大模型(MLLMs)的发展,该任务受到广泛关注。为追求更详细准确的描述,近期研究越来越多地采用强化学习(RL)。然而,现有基于RL的描述方法及评估指标往往聚焦单一质量维度,导致核心维度间的权衡。例如,以实用性为导向的目标可能诱发冗长、幻觉或噪声描述,虽提升下游问答表现,却损害流畅性;而竞技场式目标则倾向于流畅但泛化的描述,实用性不足。为此,我们提出一种更均衡的强化学习框架,联合优化实用性感知的正确性、参考覆盖度和语言质量。为有效优化连续多目标奖励,我们采用类GDPO的奖励解耦归一化策略,并引入长度条件奖励掩码,实现更适合图像描述的长度惩罚。在LLaVA-1.5-7B和Qwen2.5-VL 3B、7B基线模型上,本方法持续提升描述质量,峰值提升达+13.6 DCScore、+9.0 CaptionQA和+29.0 CapArena。

原文摘要 · Abstract (English)

Image captioning is one of the most fundamental tasks in computer vision. Owing to its open-ended nature, it has received significant attention in the era of multimodal large language models (MLLMs). In pursuit of ever more detailed and accurate captions, recent work has increasingly turned to reinforcement learning (RL). However, existing captioning-RL methods and evaluation metrics often emphasize a narrow notion of caption quality, inducing trade-offs across core dimensions of captioning. For example, utility-oriented objectives can encourage noisy, hallucinated, or overlong captions that improve downstream question answering while harming fluency, whereas arena-style objectives can favor fluent but generic descriptions with limited usefulness. To address this, we propose a more balanced RL framework that jointly optimizes utility-aware correctness, reference coverage, and linguistic quality. In order to effectively optimize the resulting continuous multi-objective reward formulation, we apply GDPO-style reward-decoupled normalization to continuous-valued captioning rewards and show that it improves performance over vanilla GRPO. Additionally, we introduce length-conditional reward masking, yielding a more suitable length penalty for captioning. Across LLaVA-1.5-7B and Qwen2.5-VL 3B and 7B base models, our method consistently improves caption quality, with peak gains of +13.6 DCScore, +9.0 CaptionQA, and +29.0 CapArena across different models.

图像描述强化学习多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。