提出在线视频密集描述生成新方法,实时输出精准时序描述
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
- 采用分因子自回归解码,逐段建模视觉特征并利用前序上下文
- 相比现有方法减少20%计算量,支持长视频处理且描述更频繁全面
- 适合需要实时视频理解与自动标注的场景
为视频生成准确、密集的描述仍是研究难点。当前多数模型需一次性处理完整视频。本文提出一种高效在线方法,在无未来帧信息下持续输出频繁、详细且时序对齐的描述。模型采用新型自回归分因子解码架构,对每个时间片段的视觉特征序列建模,输出局部化描述,并有效利用前序视频片段的上下文信息。该机制使模型能根据实际内容生成更丰富、更频繁的描述,而非模仿训练数据。此外,我们设计了高效的训练与推理优化策略,支持处理更长视频。实验表明,该方法在性能上优于离线与在线基线,计算量减少20%。生成的标注更加全面,可应用于自动视频打标与大规模视频数据采集。
原文摘要 · Abstract (English)
Generating automatic dense captions for videos that accurately describe their contents remains a challenging area of research. Most current models require processing the entire video at once. Instead, we propose an efficient, online approach which outputs frequent, detailed and temporally aligned captions, without access to future frames. Our model uses a novel autoregressive factorized decoding architecture, which models the sequence of visual features for each time segment, outputting localized descriptions and efficiently leverages the context from the previous video segments. This allows the model to output frequent, detailed captions to more comprehensively describe the video, according to its actual local content, rather than mimic the training data. Second, we propose an optimization for efficient training and inference, which enables scaling to longer videos. Our approach shows excellent performance compared to both offline and online methods, and uses 20\% less compute. The annotations produced are much more comprehensive and frequent, and can further be utilized in automatic video tagging and in large-scale video data harvesting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。