arXiv:2603.11439cs.CV2026-03中稿 · CVPR

分离定位与描述任务,提升视频字幕生成精度

Stay in your Lane: Role Specific Queries with Overlap Suppression Loss for Dense Video Captioning

  • 用角色专属查询分离定位与描述任务
  • 通过重叠抑制损失减少时间重叠,提高定位精确度
  • 轻量模块增强语义丰富性,适合多事件视频理解

密集视频字幕(DVC)是一项挑战性的多模态任务,需在视频中时序定位多个事件并用自然语言描述。现有基于查询的框架因共享查询导致定位与描述任务间干扰严重,且存在时间冗余问题。本文提出角色专属查询机制,将定位与描述拆分为独立组件,使各任务专注自身职责;通过对比对齐保证对应输出语义一致,实现行为协同。同时设计新颖的抑制机制,惩罚查询间的时序重叠,促使模型学习互不重叠的事件区域,提升定位精度。此外引入轻量级模块,捕捉核心事件概念,以概念级表示增强字幕语义丰富性。在YouCook2和ActivityNet Captions两大主流基准上实验验证了方法有效性。

原文摘要 · Abstract (English)

Dense Video Captioning (DVC) is a challenging multimodal task that involves temporally localizing multiple events within a video and describing them with natural language. While query-based frameworks enable the simultaneous, end-to-end processing of localization and captioning, their reliance on shared queries often leads to significant multi-task interference between the two tasks, as well as temporal redundancy in localization. In this paper, we propose utilizing role-specific queries that separate localization and captioning into independent components, allowing each to exclusively learn its role. We then employ contrastive alignment to enforce semantic consistency between the corresponding outputs, ensuring coherent behavior across the separated queries. Furthermore, we design a novel suppression mechanism in which mutual temporal overlaps across queries are penalized to tackle temporal redundancy, supervising the model to learn distinct, non-overlapping event regions for more precise localization. Additionally, we introduce a lightweight module that captures core event concepts to further enhance semantic richness in captions through concept-level representations. We demonstrate the effectiveness of our method through extensive experiments on major DVC benchmarks YouCook2 and ActivityNet Captions.

视频字幕多任务学习时间定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。