arXiv:2607.28509cs.CV2026-07

让视频描述精准关联多个参考图,提升生成内容真实性

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

论文配图:RefCaptioner: Multi-Reference Image-Grounded Video Captioning
图 1 · 摘自论文原文
  • 两阶段微调框架融合混合数据SFT与分层覆盖率折扣强化学习
  • 在2万段视频上实现跨参考一致性与干扰项剔除,准确率显著提升
  • 适合需要高精度视觉定位的视频理解与生成任务

现有视频字幕模型虽能生成自然描述,但无法显式将局部视觉元素与多张参考图像对齐。本文提出多参考图像引导的视频字幕新任务,要求在短语级别实现参考图像的精确绑定,并构建了包含20,000段视频和171,354张参考图像的语料库。提出RefCaptioner,通过混合数据SFT与分层覆盖率折扣强化学习(Hierarchical Coverage-Discounted GRPO),联合优化参考选择、短语级绑定、干扰项剔除及跨参考一致性,同时保持通用视频字幕能力。进一步构建MRVBench基准,用于评估真实世界与AI生成视频上的事实性与多参考对齐能力。实验表明,RefCaptioner在开源模型中表现最佳,且在标准视频字幕基准上仍具竞争力。人工评估确认其字幕更受标注者青睐,可实现更忠实于源视频的重建,适用于开源与专有视频生成器。

原文摘要 · Abstract (English)

Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference grounding, and propose RefCaptioner, a two-stage post-training framework for this task. RefCaptioner combines mixed-data SFT with Hierarchical Coverage-Discounted GRPO to jointly improve reference selection, phrase-level binding, distractor rejection, and cross-reference consistency while preserving general video-captioning ability. To support training, we construct a corpus containing $20,000$ videos and 171,354 reference images. We further introduce MRVBench, a benchmark for evaluating caption factuality and multi-reference grounding on both real-world and AI-generated videos. Experiments show that RefCaptioner achieves the best overall performance among the open-source models while remaining competitive on standard video captioning benchmarks. Human evaluation further confirms that its captions are preferred by annotators and enable more source-faithful video reconstruction with both open-source and proprietary video generators.

视频生成多参考对齐字幕生成事实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。