提出音视频联合编辑的视频检索新任务,支持跨模态修改查询。
CoVA: Text-Guided Composed Video Retrieval for Audio-Visual Content
- 设计音视频联合编辑检索任务,融合视觉与听觉变化描述。
- 构建包含跨模态差异的AV-Comp数据集,支持多模态编辑验证。
- 提出AVT模型,按需对齐文本与最相关模态,提升检索精度。
组成视频检索(CoVR)旨在通过参考视频和指定视觉修改的文本查询,从大规模视频库中检索目标视频。然而,现有基准仅关注视觉变化,忽略了在视觉相似但音频不同的情况下检索的需求。为此,我们提出音视频联合组成的视频检索任务(CoVA),同时考虑视觉与听觉变化。为支持该任务,我们构建了AV-Comp数据集,包含具有跨模态变化的视频对及其对应的文本查询。我们还提出了音视频文本组合融合(AVT)方法,通过有选择地对齐查询与最相关模态来整合视频、音频和文本特征。AVT在性能上优于传统单模态融合方法,可作为CoVA任务的强基线。更多示例可访问 https://perceptualai-lab.github.io/CoVA/。
原文摘要 · Abstract (English)
Composed Video Retrieval (CoVR) aims to retrieve a target video from a large gallery using a reference video and a textual query specifying visual modifications. However, existing benchmarks consider only visual changes, ignoring videos that differ in audio despite visual similarity. To address this limitation, we introduce Composed retrieval for Video with its Audio CoVA, a new retrieval task that accounts for both visual and auditory variations. To support this, we construct AV-Comp, a benchmark consisting of video pairs with cross-modal changes and corresponding textual queries that describe the differences. We also propose AVT Compositional Fusion (AVT), which integrates video, audio, and text features by selectively aligning the query to the most relevant modality. AVT outperforms traditional unimodal fusion and serves as a strong baseline for CoVA. Examples from the proposed dataset, including both visual and auditory information, are available at https://perceptualai-lab.github.io/CoVA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。