提出显式时序语义建模框架,提升视频密集字幕生成精度
Explicit Temporal-Semantic Modeling for Dense Video Captioning via Context-Aware Cross-Modal Interaction
- 通过跨模态检索聚合相关帧,生成时序连贯的文本特征
- 在ActivityNet和YouCook2上达到当前最佳性能,提升显著
- 适合关注视频事件理解与多模态融合的研究者
密集视频字幕任务需同时定位并描述未剪辑视频中的关键事件。现有方法多依赖附加先验知识与复杂多任务架构,但其隐式建模方式仅使用帧级或碎片化视频特征,难以捕捉事件序列间的时序一致性及视觉上下文的整体语义。为此,本文提出显式时序-语义建模框架CACMI,融合视频内在时序特性与文本语料语言语义。模型包含两个核心模块:跨模态帧聚合通过跨模态检索整合相关帧,提取时序一致、事件对齐的文本特征;上下文感知特征增强利用查询引导注意力,融合视觉动态与伪事件语义。在ActivityNet Captions与YouCook2数据集上的大量实验表明,CACMI在密集视频字幕任务中达到当前最优性能。
原文摘要 · Abstract (English)
Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competitive performance. However, these pipelines rely on implicit modeling that uses frame-level or fragmented video features, failing to capture the temporal coherence across event sequences and comprehensive semantics within visual contexts. To address this, we propose an explicit temporal-semantic modeling framework called Context-Aware Cross-Modal Interaction (CACMI), which leverages both latent temporal characteristics within videos and linguistic semantics from text corpus. Specifically, our model consists of two core components: Cross-modal Frame Aggregation aggregates relevant frames to extract temporally coherent, event-aligned textual features through cross-modal retrieval; and Context-aware Feature Enhancement utilizes query-guided attention to integrate visual dynamics with pseudo-event semantics. Extensive experiments on the ActivityNet Captions and YouCook2 datasets demonstrate that CACMI achieves the state-of-the-art performance on dense video captioning task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。