用文本-动作对比损失提升视频记忆度预测,助力更客观的视频摘要。
Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization
- 通过文本相似性构建动作正负样本对,增强运动特征表示。
- 在两个数据集上达到当前最佳性能,记忆度预测更准确。
- 可应用于视频摘要,减少人工标签主观性,适合内容生成研究者。
视频记忆度指观众观看后回忆视频的能力,在打造持久印象的内容中至关重要。现有模型虽提取多模态特征预测记忆度,但常忽略运动线索;运动特征提取器在微调阶段因缺乏标注数据导致表征退化。本文提出文本-运动跨模态对比损失(TMCCL),通过视频间文本描述相似性构建目标动作的正负样本集,使语义相关动作拥有相似特征表示,从而提升记忆度预测精度。该模型在两个视频记忆度预测数据集上达领先表现。此外,记忆度预测的应用潜力尚未充分探索。为此,我们提出记忆度加权修正视频摘要方法(MWCVS),利用记忆度预测降低视频摘要标签的主观性。在两个摘要数据集上的实验验证了其有效性,展示了记忆度预测在实际任务中的前景。
原文摘要 · Abstract (English)
Video memorability refers to the ability of videos to be recalled after viewing, playing a crucial role in creating content that remains memorable. Existing models typically focus on extracting multimodal features to predict video memorability scores but often fail to fully utilize motion cues. The representation of motion features is compromised during the fine-tuning phase of the motion feature extractor due to a lack of labeled data. In this paper, we introduce the Text-Motion Cross-modal Contrastive Loss (TMCCL), a multimodal video memorability prediction model designed to enhance the representation of motion features. We tackle the challenge of improving motion feature representation by leveraging text description similarities across videos to establish positive and negative motion sample sets for a given target. This enhancement allows the model to learn similar feature representations for semantically related motion content, resulting in more accurate memorability predictions. Our model achieves state-of-the-art performance on two video memorability prediction datasets. Moreover, the potential applications of video memorability prediction have been underexplored. To address this gap, we present Memorability Weighted Correction for Video Summarization (MWCVS), using video memorability prediction to reduce subjectivity in video summarization labels. Experimental results on two video summarization datasets demonstrate the effectiveness of MWCVS, showcasing the promising applications of video memorability prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。