用视觉事实差异做奖励,让图像描述更准又不漏关键信息。
ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison

- 以原子级视觉事实差异为奖励单位,精细定位错误类型
- 在160张图的诊断集上显著平衡了幻觉与遗漏问题
- 适合需要精准描述的场景,如医疗、自动驾驶领域
长序列图像描述在强化学习中面临奖励粒度粗的问题:整体序列评分掩盖了局部事实性错误。好的密集描述应既准确又全面,避免幻觉且不遗漏显著细节。但成对偏好、参考指标和整体标量奖励将局部错误压缩为单一序列信号,模糊了真实性和覆盖率的权衡。我们提出ClaimDiff-RL,利用参考条件下的原子事实差异作为奖励单元。给定图像、生成描述和参考描述,多模态判别器枚举视觉相关的差异,验证每项差异是否符合图像,标注开放词汇的错误类型与严重程度,并生成逐差统计用于奖励构建。这使得幻觉和遗漏显著事实可分别测量与调节。实验表明,整体标量奖励会以增加遗漏为代价减少幻觉,而ClaimDiff-RL揭示了这一权衡并实现更平衡的运行点。在160张人工标注的诊断基准、公开描述数据集及VQA任务上,ClaimDiff-RL提升了幻觉与遗漏的平衡性,保持通用能力,甚至在对象计数、空间关系、场景识别等细粒度维度上超越Gemini-3-Pro-Preview。结果表明,可分类、可验证的事实差异是实现精细、可诊断描述强化学习的有效奖励单元。
原文摘要 · Abstract (English)
Long-form image captioning exposes a reward granularity problem in RL: captions are judged as whole sequences, while the important errors occur at the level of individual visual claims. A good dense caption should be both faithful and informative, avoiding hallucination without omitting salient details. Yet pairwise preferences, reference-based metrics, and holistic scalar rewards compress these local errors into a single sequence-level signal, obscuring the tradeoff between factuality and coverage. We introduce ClaimDiff-RL, a framework that uses reference-conditioned atomic claim differences as the reward unit for caption RL. Given an image, an actor caption, and a reference caption, a multimodal judge enumerates visually grounded differences, verifies each difference against the image, assigns open-vocabulary error types and severity levels, and produces per-difference statistics for reward composition. This makes hallucinated claims and omitted salient facts separately measurable and tunable. Experiments show that holistic scalar rewards can reduce hallucination by increasing missing facts, while ClaimDiff-RL exposes this faithfulness and coverage tradeoff and enables more balanced operating points. On a 160-image human-labeled diagnostic benchmark, public captioning benchmarks, and VQA benchmarks, ClaimDiff-RL improves the hallucination--missing-fact balance, preserves general capability, and even surpasses Gemini-3-Pro-Preview on several fine-grained Capability dimensions such as object counting, spatial relations, and scene recognition. These results suggest that typed, verifiable claim differences are an effective reward unit for fine-grained and diagnosable caption RL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。