通过递归关联多模态编码器,提升视频理解中复杂动作与场景的捕捉能力。
Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
- 将预训练编码器视为'超神经元',通过递归融合多模态信息
- 在像素级追踪中平均交并比提升2.7%,时间连贯性下降8.8%
- 适用于视频追踪、识别、对话与编辑,尤其适合复杂动态场景
视频理解被视为通向世界建模的关键步骤,是人工智能长期研究的重要方向。近期,多模态基础模型通过大规模预训练展现出巨大潜力,有效利用对比学习对齐不同模态的编码器。为进一步提升在复杂目标运动和多样化视频场景下的性能,我们提出通过更深层的多模态交互增强这一对齐机制,以更好理解复杂运动与多样场景。为此,我们构建统一的超编码网络(Super Encoding Network, SEN),通过基础模型中多模态编码器的递归关联实现此类深度交互。具体而言,我们将已训练好的编码器创造性地视为‘超神经元’,设计递归关联(RA)模块,以知识集成、分发与提示的方式,递归地融合多模态信息。该方法能有效编码深层次多模态交互,用于下游视频理解任务。大量实验表明,SEN显著提升了四项代表性视频任务:追踪、识别、对话与编辑。例如,在像素级追踪中,平均交并比(Jaccard)提升2.7%,时间连贯性(TC)下降8.8%;在单次视频编辑中,文本对齐度提升6.4%,帧一致性提高4.1%,均优于主流方法如CaDeX++与Tune-A-Video。
原文摘要 · Abstract (English)
Video understanding has been considered as one critical step towards world modeling, which is an important long-term problem in AI research. Recently, multimodal foundation models have shown such potential via large-scale pretraining. These models effectively align encoders of different modalities via contrastive learning. To further enhance performance on complex target movements and diversified video scenes, we propose to augment this alignment with deeper multimodal interactions, which are critical for understanding complex target movements with diversified video scenes. To fill this gap, we propose a unified Super Encoding Network (SEN) for video understanding, which builds up such distinct interactions through the recursive association of multimodal encoders in the foundation models. Specifically, we creatively treat those well-trained encoders as ``super neurons" in our SEN. Via designing a Recursive Association (RA) block, we progressively fuse multi-modalities with the input video, based on knowledge integrating, distributing, and prompting of super neurons in a recursive manner. In this way, our SEN can effectively encode deeper multimodal interactions for prompting various video understanding tasks in the downstream. Extensive experiments show that our SEN can remarkably boost the four most representative video tasks, including tracking, recognition, chatting, and editing, e.g., for pixel-level tracking, the average jaccard index improves 2.7%, and temporal coherence(TC) drops by 8.8% compared to the popular CaDeX++ approach. For one-shot video editing, textual alignment improves 6.4%, and frame consistency increases by 4.1% compared to the Tune-A-Video approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。