让重复视频有唯一描述,提升精准检索能力
It's Just Another Day: Unique Video Captioning by Discriminative Prompting
- 通过区分性提示预测可区分属性,生成唯一标题
- 在第一人称视频上文本检索准确率提升15%
- 适用于重复动作多的视频场景,如自拍或循环拍摄
长视频中常出现重复的动作、事件和镜头,这些片段往往被赋予相同的字幕,导致通过文本搜索难以精确定位目标片段。本文提出唯一字幕生成问题:针对具有相同字幕的多个片段,为每个片段生成能唯一标识它的新字幕。我们提出区分性提示字幕生成(CDP)方法,通过预测可区分的属性来生成唯一性字幕。构建了两个新的唯一字幕基准数据集,分别基于第一人称视角视频和时间循环影片(timeloop movies),二者均存在大量重复动作。实验表明,使用CDP生成的字幕使第一人称视频的文本到视频检索R@1提升15%,在时间循环影片中提升10%。
原文摘要 · Abstract (English)
Long videos contain many repeating actions, events and shots. These repetitions are frequently given identical captions, which makes it difficult to retrieve the exact desired clip using a text search. In this paper, we formulate the problem of unique captioning: Given multiple clips with the same caption, we generate a new caption for each clip that uniquely identifies it. We propose Captioning by Discriminative Prompting (CDP), which predicts a property that can separate identically captioned clips, and use it to generate unique captions. We introduce two benchmarks for unique captioning, based on egocentric footage and timeloop movies - where repeating actions are common. We demonstrate that captions generated by CDP improve text-to-video R@1 by 15% for egocentric videos and 10% in timeloop movies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。