arXiv:2605.28023cs.CVcs.AI2026-05被引 1

用视觉信号验证文本一致性,提升图文生成的准确性和可靠性。

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning

论文配图:VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
图 1 · 摘自论文原文
  • 引入见证-裁决奖励机制,结合参考文本与视觉信号验证事实一致性。
  • 80亿参数模型在多任务上超越开源与闭源最佳模型,人类评估更贴合事实。
  • 适用于弱到强的强化学习训练,可推广至跨任务、跨模态场景。

视觉描述生成要求模型忠实捕捉视觉内容,同时避免遗漏与幻觉。当前主流的多模态大模型(MLLM)通过规模扩展和高质量数据已取得优异表现。近年来,强化学习(RL)成为提升精度与覆盖范围的关键路径,但现有图像描述奖励设计缺乏细粒度、可靠的真伪验证信号,限制了其效果。为此,本文提出VCap——一种基于见证-裁决机制的奖励方法:将参考描述(见证)与视觉信号(裁决)配对,显式验证参考描述与策略生成描述在视觉基础上的事实一致性。该机制提供符合超几何分布精度的奖励信号,实现对生成质量的精准评估。即使参考文本存在缺陷,也能有效驱动模型学习,支持从弱监督到强性能的强化学习泛化。实验表明,使用VCap训练的80亿参数模型,在多个图像与视频描述基准上超越开源及闭源最新模型;人工评估进一步证实其与真实一致性的高度对齐。此外,VCap显著提升了模型感知能力,具备跨任务泛化性,优于最优-多抽样蒸馏(best-of-N distillation),挑战了关于强化学习视觉推理(RLVR)的既有认知。

原文摘要 · Abstract (English)

Visual captioning requires models to capture visual content faithfully while minimizing both omission and hallucination. As the dominant paradigm for captioning, MLLMs have achieved strong performance through scaling and high-quality data. Recently, RL has emerged as a key route to driving MLLMs toward higher precision and broader coverage, however, existing reward designs for captioning fail to provide fine-grained and reliable signals for factual verification, limiting their effectiveness. To address this, we propose VCap, a Witness-Adjudicator reward that pairs the reference caption (a witness) with the visual signal (an adjudicator). By explicitly verifying factual consistency between the reference and policy-generated captions grounded in the visual signal, VCap delivers a reward signal with hypergeometric-distribution-level precision for caption quality verification. This design enables effective learning even from imperfect references, facilitating weak-to-strong generalization in RL training. In our experiments, an 8B model trained with VCap outperforms open- and closed-source SOTA models on multiple image and video captioning benchmarks. Human evaluation further confirms its strong alignment with factual correctness. Additionally, VCap improves MLLM perceptual capability, generalizes across tasks, and surpasses best-of-N distillation, challenging prior assumptions about RLVR.

视觉描述强化学习事实验证多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。