arXiv:2605.06229cs.CV2026-05被引 1

让视频搜索关注被忽略的背景区域,提升复杂场景检索效果

Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search

论文配图:Look Beyond Saliency: Low-Attention Guided Dual Encoding for Video Semantic Search
图 1 · 摘自论文原文
  • 用反向注意力机制捕捉被忽视的背景信息
  • 在拥挤场景下召回率显著优于现有方法
  • 无需额外训练,可直接融合现有模型

密集场景下的视频语义搜索仍具挑战性,因视觉编码器倾向于关注显著的前景区域,而忽略具有语义重要性的背景区域。本文提出逆向注意力嵌入机制,显式捕捉并突出这些被忽略的区域。通过将逆向注意力嵌入与传统视觉嵌入结合,该方法在不增加训练成本的前提下显著提升语义检索性能。初步实验与消融研究显示,在拥挤环境下的视频语义搜索中,召回率相比现有方法有明显提升。

原文摘要 · Abstract (English)

Video semantic search in densely crowded scenes remains a challenging task due to visual encoders tendency to prioritize salient foreground regions while neglecting contextually important, background areas. We propose an Inverse Attention Embedding mechanism that explicitly captures and highlights these overlooked regions. By combining inverse attention embeddings with traditional visual embeddings, our method significantly enhances semantic retrieval performance without additional training. Initial experiments and ablation studies demonstrate promising improvements over existing approaches in recall for video semantic search in crowded environments.

视频检索注意力机制语义搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。