arXiv:2607.21022cs.CV2026-07

通过迭代修正提升视频描述完整性,不重训练模型也能减少幻觉。

ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning

论文配图:ProCap: Prominence-guided Object Rectification for Faithful and Comprehensive Video Captioning
图 1 · 摘自论文原文
  • 根据显著性、持续性和关系动态排序物体,逐轮注入关键信息。
  • 人类评估显示完整性提升48%,幻觉减少45%,无需参考文本。
  • 适合追求高效、可靠视频描述的无障碍与检索应用。

提升视频描述质量通常需要重新训练大型视觉语言模型,成本高且不实用。现有无训练替代方法虽基于检测到的物体来约束描述以减少幻觉,但仅执行一次固定修正,未优先处理重要物体,导致语义关键内容遗漏。我们提出一种显著性感知、迭代式后处理修正框架,无需修改底层模型参数:轻量级评分机制依据空间显著性、时间持续性和关系动态对检测物体排序;迭代式提示驱动修正环路利用该排序,多轮逐步将缺失但上下文相关的物体注入描述中。在MSVD和MSR-VTT上验证,使用基于物体的自动指标、110人参与的人类研究及与ChatGPT和Gemini的定性对比;在人类评估中,该框架使感知完整性最高提升48%,幻觉最多降低45%,且无需重训练或参考字幕。结果表明,显著性引导的迭代修正是一种轻量、可扩展、模型无关的路径,能生成更完整、可信的视频描述,直接适用于无障碍、信息检索等多媒体理解场景。

原文摘要 · Abstract (English)

Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing training-free alternatives instead ground captions in detected objects to curb hallucination, but apply only a single, fixed correction pass without prioritizing which objects matter most, leaving semantically significant content omitted. We propose a prominence-aware, iterative post-hoc rectification framework that overcomes both limitations without modifying the underlying captioning model's parameters: a lightweight scoring mechanism ranks detected objects by spatial saliency, temporal persistence, and relational dynamics, and an iterative, prompt-driven refinement loop uses this ranking to progressively inject missing yet contextually relevant objects into the caption over multiple rounds. We validate the framework on MSVD and MSR-VTT using object-grounded automatic metrics, a 110-participant human study, and qualitative comparison against ChatGPT and Gemini; in human evaluation, the framework raises perceived completeness by up to 48% and reduces hallucination by up to 45% relative to a strong pretrained captioning baseline, all without retraining or reference captions. These results position prominence-guided iterative rectification as a lightweight, scalable, and model-agnostic route to more complete and trustworthy video captioning, with direct relevance to accessibility, retrieval, and other multimedia understanding applications.

视频描述去幻觉迭代修正轻量级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。