arXiv:2508.02340cs.CVcs.IR2025-08中稿 · ACMMM2025

让视频搜索更全面:通过多空间解耦提升跨场景检索能力

Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search

  • 为每类特征学习独立的语义空间,避免信息混淆
  • 在TRECVID 2016-2023上平均提升mAP达3.2个百分点
  • 适合需要高召回与多样化结果的视频检索场景

即兴视频搜索(AVS)需用文本查询从大量未标注短视频中找出多个相关片段。其主要挑战在于相关视频的视觉多样性——例如‘男女在室内共舞’可能涵盖亮堂厅堂、昏暗酒吧或黑白动画等不同场景。现有方法通常将多特征融合至单一语义空间,忽略了多样性的需求。为此,本文提出LPD(Learning Partially Decorrelated common spaces),核心包含两个创新:特征专属的公共空间构建与去相关损失。LPD为每个视频和文本特征分别学习独立的公共空间,并通过去相关损失使不同空间中的负样本排序差异最大化。为进一步保障多空间收敛的一致性,设计了基于熵的公平多空间三元组排序损失。在TRECVID AVS基准(2016–2023)上的实验充分验证了LPD的有效性,且可视化显示其显著增强了检索结果的多样性。

原文摘要 · Abstract (English)

Ad-hoc Video Search (AVS) involves using a textual query to search for multiple relevant videos in a large collection of unlabeled short videos. The main challenge of AVS is the visual diversity of relevant videos. A simple query such as "Find shots of a man and a woman dancing together indoors" can span a multitude of environments, from brightly lit halls and shadowy bars to dance scenes in black-and-white animations. It is therefore essential to retrieve relevant videos as comprehensively as possible. Current solutions for the AVS task primarily fuse multiple features into one or more common spaces, yet overlook the need for diverse spaces. To fully exploit the expressive capability of individual features, we propose LPD, short for Learning Partially Decorrelated common spaces. LPD incorporates two key innovations: feature-specific common space construction and the de-correlation loss. Specifically, LPD learns a separate common space for each video and text feature, and employs de-correlation loss to diversify the ordering of negative samples across different spaces. To enhance the consistency of multi-space convergence, we designed an entropy-based fair multi-space triplet ranking loss. Extensive experiments on the TRECVID AVS benchmarks (2016-2023) justify the effectiveness of LPD. Moreover, diversity visualizations of LPD's spaces highlight its ability to enhance result diversity.

视频搜索多模态去相关语义空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。