arXiv:2512.13015cs.CV2025-12被引 2

用统一模型预测视频下一场景,推动视觉系统理解时间因果关系。

What Happens Next? Next Scene Prediction with a Unified Video Model

  • 融合双模型,通过潜变量桥接实现上下文到未来场景的推理。
  • 在自建数据集上三阶段训练,达成当前最佳预测效果。
  • 适合研究视频理解与生成的开发者,尤其关注时序推理者。

近期统一多模态模型在视觉生成方面取得显著进展,但其主要聚焦于文本到视频等常规任务,对统一模型的时间推理潜力挖掘不足。为弥补这一缺口,我们提出「下一场景预测」(Next Scene Prediction, NSP)新任务,推动统一视频模型向时间与因果推理发展。不同于文本到视频生成,NSP需基于前序上下文预测合理未来,要求更深层次的理解与推理。为此,我们构建统一框架:以Qwen-VL进行理解,用LTX实现生成,二者通过潜变量嵌入与连接模块协同。模型在自建大规模NSP数据集上分三阶段训练:文本到视频预训练、监督微调及基于因果一致性奖励的强化学习(GRPO)。实验表明,该模型在基准测试中达到当前最优性能,显著提升通用多模态系统对未来事件的预判能力。

原文摘要 · Abstract (English)

Recent unified models for joint understanding and generation have significantly advanced visual generation capabilities. However, their focus on conventional tasks like text-to-video generation has left the temporal reasoning potential of unified models largely underexplored. To address this gap, we introduce Next Scene Prediction (NSP), a new task that pushes unified video models toward temporal and causal reasoning. Unlike text-to-video generation, NSP requires predicting plausible futures from preceding context, demanding deeper understanding and reasoning. To tackle this task, we propose a unified framework combining Qwen-VL for comprehension and LTX for synthesis, bridged by a latent query embedding and a connector module. This model is trained in three stages on our newly curated, large-scale NSP dataset: text-to-video pre-training, supervised fine-tuning, and reinforcement learning (via GRPO) with our proposed causal consistency reward. Experiments demonstrate our model achieves state-of-the-art performance on our benchmark, advancing the capability of generalist multimodal systems to anticipate what happens next.

视频预测统一模型时序推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。