arXiv:2606.31326cs.CV2026-06被引 1

统一视频理解与生成,用共享语义库和混合架构实现跨任务协同。

Bridging Video Understanding and Generation in a Unified Framework

论文配图:Bridging Video Understanding and Generation in a Unified Framework
图 1 · 摘自论文原文
  • 共享词汇表融合文本与视觉表征,结合自回归与扩散模型
  • 关键帧语义预测引导高分辨率视频生成,提升时序连贯性
  • 在生成与理解双任务上均表现优异,适合多模态研究者

近年来,统一图像生成与理解已受到广泛研究,但将此类统一建模范式拓展至视频领域仍进展有限。核心挑战在于:视频理解需要紧凑、有区分度的语义表征,而视频生成则需保留视觉细节与时间连贯性的密集信号。视频天然兼具空间语义与时间动态,比静态图像更适合作为统一多模态建模的载体。本文提出Vega,一个统一视频理解与生成的框架。Vega采用共享词汇表联合建模文本与视觉表示,并引入混合架构,结合自回归(AR)预测与基于扩散的渲染。具体而言,AR模型聚焦于关键帧的语义化视觉标记预测,提供结构化表征以指导扩散模块生成高分辨率、密集的视频帧。大量实验表明,Vega在视频生成基准VBench及视频理解基准VideoMME上均取得强劲性能。

原文摘要 · Abstract (English)

Recently, unified image generation and understanding have been extensively explored. However, extending such unified modeling paradigms to the video domain remains largely underexplored. A central challenge is that video understanding favors compact, discriminative semantic representations, whereas video generation requires dense signals that preserve visual details and temporal coherence. Videos naturally capture both spatial semantics and temporal dynamics, making them a more suitable modality for unified multimodal modeling compared to static images. In this paper, we propose Vega, a unified framework that bridges video understanding and generation. Vega leverages a shared vocabulary to jointly model text and visual representations and employs a hybrid architecture combining autoregressive (AR) prediction with diffusion-based rendering. Specifically, the AR model focuses on predicting semantically meaningful visual tokens for keyframes, providing a structured representation that guides the diffusion module in rendering dense, high-resolution video frames. Extensive experiments demonstrate that Vega achieves strong performance on video generation benchmarks such as VBench and video understanding benchmarks like VideoMME.

视频生成统一建模扩散模型自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。