arXiv:2412.02071cs.CV2024-12CVPR被引 11

让视频每帧都描述动作进展,更精准捕捉动态变化。

Progress-Aware Video Frame Captioning

  • 设计新模型ProgressCaptioner,捕捉动作序列的细粒度时序变化。
  • 在自建数据集FrameCap上表现超越现有模型,提升时序精度。
  • 适用于关键帧选取和视频理解,实用性强。

虽然图像字幕为单张图像提供孤立描述,视频字幕则为整段视频生成单一叙述,本文探索了一个重要中间地带:帧级进度感知视频字幕生成。该新任务旨在生成时间上精细的字幕,不仅准确描述每一帧,还捕捉动作在整个视频序列中的细微演变。尽管现有领先视觉语言模型能力强大,但往往难以察觉帧间差异的微妙之处。为此,我们提出ProgressCaptioner模型,专门用于捕捉动作序列中的细粒度时序动态。同时,我们构建了FrameCap数据集以支持训练,并建立了FrameCapEval评估基准以衡量字幕质量。实验结果表明,ProgressCaptioner显著优于现有主流字幕模型,生成的字幕能精确反映动作进展,树立了视频字幕时序精度的新标准。最后,我们展示了该方法在辅助关键帧选择和推动视频理解方面的实际应用,凸显其广泛适用性。

原文摘要 · Abstract (English)

While image captioning provides isolated descriptions for individual images, and video captioning offers one single narrative for an entire video clip, our work explores an important middle ground: progress-aware video captioning at the frame level. This novel task aims to generate temporally fine-grained captions that not only accurately describe each frame but also capture the subtle progression of actions throughout a video sequence. Despite the strong capabilities of existing leading vision language models, they often struggle to discern the nuances of frame-wise differences. To address this, we propose ProgressCaptioner, a captioning model designed to capture the fine-grained temporal dynamics within an action sequence. Alongside, we develop the FrameCap dataset to support training and the FrameCapEval benchmark to assess caption quality. The results demonstrate that ProgressCaptioner significantly surpasses leading captioning models, producing precise captions that accurately capture action progression and set a new standard for temporal precision in video captioning. Finally, we showcase practical applications of our approach, specifically in aiding keyframe selection and advancing video understanding, highlighting its broad utility.

视频字幕动作进展时序建模多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。