arXiv:2607.05093cs.CV2026-07中稿 · AAAI

用多尺度卷积与动态路由提升视频文本定位精度

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing

论文配图:Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing
图 1 · 摘自论文原文
  • 采用多尺度卷积捕捉不同时间粒度的运动模式
  • 动态路由机制实现复杂多对多对应关系建模,达42.9% [email protected]
  • 适合需要细粒度视频语言对齐的研究与应用

视频-文本时间定位需精准对齐自然语言查询与对应视频片段,是多模态理解的核心挑战。本文提出新框架,解决现有方法在层次化时间结构建模不足及难以处理跨模态复杂多对多对应关系的问题。引入多尺度时间卷积编码器,捕捉从帧间瞬时变化到长动作序列的不同时间粒度的运动模式;进一步提出基于胶囊的动态路由机制,通过结构化一致更新迭代优化片段-查询关联,支持非单调对齐的灵活建模。二者通过多任务学习目标统一,联合优化时间边界回归、跨模态语义对齐与胶囊多样性。在ActivityNet Captions上实验表明,模型取得42.9% [email protected]和41.1%平均IoU,超越强基线的Transformer模型,同时保持高效计算。结果验证:结合层次化时间建模与结构化语义路由,可有效提升细粒度视频-语言理解能力。

原文摘要 · Abstract (English)

Video-text temporal localization requires precise alignment between natural language queries and corresponding video segments, a fundamental challenge in multimodal understanding. We present a novel framework that addresses two critical limitations of existing methods: inadequate modeling of hierarchical temporal structure and inability to handle complex many-to-many correspondences between modalities. Our approach introduces a multi-scale temporal convolutional encoder that captures motion patterns across different temporal granularities - from instantaneous frame transitions to extended action sequences. We further propose a capsule-based dynamic routing mechanism that iteratively refines segment-query associations through structured agreement updates, enabling flexible modeling of non-monotonic alignments. These components are unified through a multi-task learning objective that jointly optimizes temporal boundary regression, cross-modal semantic alignment, and capsule diversity. Extensive experiments on ActivityNet Captions demonstrate significant improvements, achieving 42.9% [email protected] and 41.1% mean IoU, surpassing strong transformer-based baselines while maintaining computational efficiency. Our results validate that combining hierarchical temporal modeling with structured semantic routing provides an effective solution for fine-grained video-language understanding.

视频定位多模态动态路由时间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。