arXiv:2510.07720cs.IR2025-10中稿 · International ACM …被引 28

通过聚类文本查询提升视频搜索语义理解能力

Queries Are Not Alone: Clustering Text Embeddings for Video Search

  • 将相似查询聚类,扩展单个查询的语义范围
  • 在5个公开数据集上优于当前最优模型
  • 适合需要精准语义匹配的视频检索场景

跨平台视频内容的快速增长凸显了先进视频检索系统的需求。传统方法依赖文本查询与视频元数据的直接匹配,难以弥合文本描述与视频内容多维特征之间的语义鸿沟。本文提出Video-Text Cluster(VTC)框架,通过聚类文本查询来捕捉更广泛的语义范围。我们设计了一种独特的聚类机制,将相关查询分组,使系统能考虑每个查询的多重含义与细微差别。该聚类过程由创新的Sweeper模块优化,用于识别并消除集群内的噪声。此外,引入Video-Text Cluster-Attention(VTC-Att)机制,根据视频内容动态调整对集群内文本特征的关注度,确保检索过程聚焦最相关的语义特征。进一步实验表明,所提模型在五个公开数据集上均超越现有最先进模型。

原文摘要 · Abstract (English)

The rapid proliferation of video content across various platforms has highlighted the urgent need for advanced video retrieval systems. Traditional methods, which primarily depend on directly matching textual queries with video metadata, often fail to bridge the semantic gap between text descriptions and the multifaceted nature of video content. This paper introduces a novel framework, the Video-Text Cluster (VTC), which enhances video retrieval by clustering text queries to capture a broader semantic scope. We propose a unique clustering mechanism that groups related queries, enabling our system to consider multiple interpretations and nuances of each query. This clustering is further refined by our innovative Sweeper module, which identifies and mitigates noise within these clusters. Additionally, we introduce the Video-Text Cluster-Attention (VTC-Att) mechanism, which dynamically adjusts focus within the clusters based on the video content, ensuring that the retrieval process emphasizes the most relevant textual features. Further experiments have demonstrated that our proposed model surpasses existing state-of-the-art models on five public datasets.

视频搜索文本聚类语义匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。