arXiv:2510.23043cs.CV2025-10被引 2

用分层锚点机制实现长视频精准时序定位

HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling

  • 分层锚点Mamba池化捕捉多粒度时序特征
  • 在Ego4D-NLQ等数据集上刷新最佳性能
  • 适合需要高精度时序定位的视频理解任务

视频时序定位旨在从无剪辑视频中定位自然语言查询的起止时间,需兼顾全局上下文与细粒度时序细节。该任务在长视频中尤为困难,现有方法常因过度下采样或依赖固定窗口而损失时序精度。我们提出HieraMamba,一种分层架构,可保留多尺度下的时序结构与语义丰富性。核心是锚点-Mamba池化(AMP)模块,利用Mamba的选通扫描生成紧凑锚点令牌,以多粒度总结视频内容。通过锚点条件与段落池化对比损失两个互补目标,促使锚点既保持局部细节又具备全局判别力。HieraMamba在Ego4D-NLQ、MAD和TACoS数据集上取得新最优结果,实现了长且无剪辑视频中的精确时序定位。

原文摘要 · Abstract (English)

Video temporal grounding, the task of localizing the start and end times of a natural language query in untrimmed video, requires capturing both global context and fine-grained temporal detail. This challenge is particularly pronounced in long videos, where existing methods often compromise temporal fidelity by over-downsampling or relying on fixed windows. We present HieraMamba, a hierarchical architecture that preserves temporal structure and semantic richness across scales. At its core are Anchor-MambaPooling (AMP) blocks, which utilize Mamba's selective scanning to produce compact anchor tokens that summarize video content at multiple granularities. Two complementary objectives, anchor-conditioned and segment-pooled contrastive losses, encourage anchors to retain local detail while remaining globally discriminative. HieraMamba sets a new state-of-the-art on Ego4D-NLQ, MAD, and TACoS, demonstrating precise, temporally faithful localization in long, untrimmed videos.

视频定位时序建模Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。