arXiv:2603.06732cs.CV2026-03

提出新框架HERO,让视频定位语言更灵活,支持没见过的词和表达。

HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos

  • 分层语义嵌入+跨模态并行优化,提升视频与语言对齐精度
  • 在新构建的两个开放词汇基准上,性能超越现有方法
  • 适合研究开放词汇视频理解、自然语言交互的学者

视频中的时间句子定位(TSGV)旨在找出与自然语言查询对应的时间片段。尽管近期取得进展,但大多数现有方法局限于封闭词汇设置,难以泛化到包含新词或多样表达的真实查询。为此,我们提出开放词汇TSGV(OV-TSGV)任务,并构建首个专用基准——Charades-OV与ActivityNet-OV,模拟真实词汇变化与同义表达。这些基准支持模型泛化能力的系统评估。为应对该任务,我们提出HERO(Hierarchical Embedding-Refinement for Open-Vocabulary grounding),一种统一框架,利用分层语言嵌入并执行并行跨模态优化。HERO联合建模多层次语义,通过语义引导的视觉过滤和对比掩码文本重构增强视频-语言对齐。在标准与开放词汇基准上的大量实验表明,HERO持续优于现有方法,尤其在开放词汇场景下表现突出,验证其强泛化能力,凸显OV-TSGV作为新研究方向的重要性。

原文摘要 · Abstract (English)

Temporal Sentence Grounding in Videos (TSGV) aims to temporally localize segments of a video that correspond to a given natural language query. Despite recent progress, most existing TSGV approaches operate under closed-vocabulary settings, limiting their ability to generalize to real-world queries involving novel or diverse linguistic expressions. To bridge this critical gap, we introduce the Open-Vocabulary TSGV (OV-TSGV) task and construct the first dedicated benchmarks--Charades-OV and ActivityNet-OV--that simulate realistic vocabulary shifts and paraphrastic variations. These benchmarks facilitate systematic evaluation of model generalization beyond seen training concepts. To tackle OV-TSGV, we propose HERO(Hierarchical Embedding-Refinement for Open-Vocabulary grounding), a unified framework that leverages hierarchical linguistic embeddings and performs parallel cross-modal refinement. HERO jointly models multi-level semantics and enhances video-language alignment via semantic-guided visual filtering and contrastive masked text refinement. Extensive experiments on both standard and open vocabulary benchmarks demonstrate that HERO consistently surpasses state-of-the-art methods, particularly under open-vocabulary scenarios, validating its strong generalization capability and underscoring the significance of OV-TSGV as a new research direction.

视频理解开放词汇跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。