arXiv:2601.16155cs.CVcs.IR2026-01中稿 · ICASSP 2026被引 14

模仿人类视觉注意力,让视频检索更精准。

HVD: Human Vision-Driven Video Representation Learning for Text-Video Retrieval

  • 分阶段筛选关键帧与视觉片段,模拟人眼聚焦过程
  • 在5个基准上达到当前最优效果,提升检索精度
  • 适合需要精准图文匹配的视频理解任务

CLIP的成功推动了文本-视频检索的发展。然而,现有方法常因文本查询稀疏导致特征交互“盲视”,难以区分关键视觉信息与背景噪声。为此,我们借鉴人类认知行为,提出人类视觉驱动(HVD)模型。该框架采用粗到细对齐机制,包含帧特征选择模块(FFSM)和块特征压缩模块(PFCM)。FFSM模拟人类宏观感知,通过选择关键帧消除时间冗余;随后PFCM通过先进注意力机制将块特征聚合为显著视觉实体,实现细粒度实体级匹配。在五个基准上的大量实验表明,HVD不仅捕捉类人视觉焦点,还实现了领先性能。

原文摘要 · Abstract (English)

The success of CLIP has driven substantial progress in text-video retrieval. However, current methods often suffer from "blind" feature interaction, where the model struggles to discern key visual information from background noise due to the sparsity of textual queries. To bridge this gap, we draw inspiration from human cognitive behavior and propose the Human Vision-Driven (HVD) model. Our framework establishes a coarse-to-fine alignment mechanism comprising two key components: the Frame Features Selection Module (FFSM) and the Patch Features Compression Module (PFCM). FFSM mimics the human macro-perception ability by selecting key frames to eliminate temporal redundancy. Subsequently, PFCM simulates micro-perception by aggregating patch features into salient visual entities through an advanced attention mechanism, enabling precise entity-level matching. Extensive experiments on five benchmarks demonstrate that HVD not only captures human-like visual focus but also achieves state-of-the-art performance.

视频检索视觉注意力CLIP多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。