arXiv:2507.17402cs.CVcs.IR2025-07ICCV被引 21

用双曲空间提升视频片段检索的层次建模能力

HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

  • 融合双曲与欧氏空间,动态融合视频特征
  • 在多个数据集上超越现有最佳方法,提升1.8%-3.2%
  • 适合需要细粒度视频语义匹配的研究者

部分相关视频检索(PRVR)旨在匹配未剪辑视频与仅描述部分内容的文本查询。现有方法因欧氏空间的几何失真,难以准确表达视频内在层次结构,忽略部分层次语义,导致时间建模不佳。为此,我们提出首个面向PRVR的双曲建模框架HLFormer,利用双曲空间学习弥补欧氏空间在层次建模上的不足。HLFormer集成洛伦兹注意力模块与欧氏注意力模块,在混合空间中编码视频嵌入,并通过均值引导自适应交互模块动态融合特征。此外,引入偏序保持损失,通过洛伦茨锥约束强化“文本 < 视频”的层级关系。该方法进一步提升了视频内容与文本查询间的部分相关性匹配效果。大量实验证明,HLFormer显著优于当前最优方法。代码已开源:https://github.com/lijun2005/ICCV25-HLFormer。

原文摘要 · Abstract (English)

Partially Relevant Video Retrieval (PRVR) addresses the critical challenge of matching untrimmed videos with text queries describing only partial content. Existing methods suffer from geometric distortion in Euclidean space that sometimes misrepresents the intrinsic hierarchical structure of videos and overlooks certain hierarchical semantics, ultimately leading to suboptimal temporal modeling. To address this issue, we propose the first hyperbolic modeling framework for PRVR, namely HLFormer, which leverages hyperbolic space learning to compensate for the suboptimal hierarchical modeling capabilities of Euclidean space. Specifically, HLFormer integrates the Lorentz Attention Block and Euclidean Attention Block to encode video embeddings in hybrid spaces, using the Mean-Guided Adaptive Interaction Module to dynamically fuse features. Additionally, we introduce a Partial Order Preservation Loss to enforce "text < video" hierarchy through Lorentzian cone constraints. This approach further enhances cross-modal matching by reinforcing partial relevance between video content and text queries. Extensive experiments show that HLFormer outperforms state-of-the-art methods. Code is released at https://github.com/lijun2005/ICCV25-HLFormer.

视频检索双曲学习跨模态匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。