arXiv:2607.05150cs.CV2026-07被引 2

用细粒度断言评分提升视频字幕生成的准确性与多样性

Claim-Level Rubric Rewards for Video Caption Reinforcement Learning

论文配图:Claim-Level Rubric Rewards for Video Caption Reinforcement Learning
图 1 · 摘自论文原文
  • 将字幕拆解为带类别标签的原子断言,逐条验证
  • 相比整体评分更少受风格干扰,避免事实错误
  • 适合需要高准确性和多样性的开放生成任务

本文提出一种名为CuRe的结构化奖励框架,旨在解决密集视频字幕生成中强化学习的奖励设计瓶颈。现有方法主要分为两类:跨异质标准的整体响应评估,或基于参考字幕的对齐评估。前者难以保证事实准确性,易受风格奖励欺骗;后者依赖严格文本对齐,无法保留开放式生成任务中的完整性和多样性。CuRe将奖励建模重构为细粒度断言级验证,通过结构化评分体系将字幕分解为类别感知的原子断言,将复杂整体评估转化为更简单可靠的断言级验证。

原文摘要 · Abstract (English)

In this paper, we introduce Claim-Level Rubric Rewards (CuRe), a structured reward framework designed to address the reward-design bottleneck in reinforcement learning for dense video captioning. Existing reward designs generally fall into two categories: holistic response-level judgment across heterogeneous criteria, or alignment-based evaluation against reference captions. However, both paradigms suffer from fundamental limitations. Holistic rewards struggle to ensure factual accuracy and are prone to stylistic reward hacking, while reference-based rewards overly rely on rigid textual alignment, failing to preserve the completeness and diversity inherent to open-ended generation tasks. To address these challenges, CuRe reformulates reward modeling as fine-grained claim-level verification. Specifically, CuRe decomposes captions into category-aware atomic claims through a structured rubric, converting holistic evaluation into simpler and more reliable claim-level verification.

视频字幕强化学习奖励设计生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。