arXiv:2603.24835cs.CV2026-03被引 1

分治式自回归框架,生成32秒长视频更稳定清晰

DCARL: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation

  • 先生成关键帧锚点,再分段自回归插值,兼顾全局结构与局部连贯
  • 在32秒长视频上实现更低的FID、FVD和相机轨迹误差
  • 适合需要高保真长时序视频生成的研究与应用

长轨迹视频生成是世界建模的关键挑战,主要受限于现有视频扩散模型(VDMs)的可扩展性。自回归模型虽可无限续写,但存在视觉漂移和控制性差的问题。为此,我们提出DCARL——一种新颖的分治式自回归框架,有效结合分治结构的稳定性与VDM的高保真生成能力。方法首先训练专用关键帧生成器,在无时间压缩的情况下建立长程全局一致的结构锚点;随后,插值生成器以重叠段自回归方式合成密集帧,利用关键帧提供全局上下文,仅依赖前一帧保持局部一致性。在大规模互联网长轨迹视频数据集上训练,该方法在视觉质量(更低的FID和FVD)与相机轨迹一致性(更低的ATE和ARE)方面均优于当前最先进的自回归及分治基线,证明其在长达32秒的视频生成中具备稳定且高保真的能力。

原文摘要 · Abstract (English)

Long-trajectory video generation is a crucial yet challenging task for world modeling primarily due to the limited scalability of existing video diffusion models (VDMs). Autoregressive models, while offering infinite rollout, suffer from visual drift and poor controllability. To address these issues, we propose DCARL, a novel divide-and-conquer, autoregressive framework that effectively combines the structural stability of the divide-and-conquer scheme with the high-fidelity generation of VDMs. Our approach first employs a dedicated Keyframe Generator trained without temporal compression to establish long-range, globally consistent structural anchors. Subsequently, an Interpolation Generator synthesizes the dense frames in an autoregressive manner with overlapping segments, utilizing the keyframes for global context and a single clean preceding frame for local coherence. Trained on a large-scale internet long trajectory video dataset, our method achieves superior performance in both visual quality (lower FID and FVD) and camera adherence (lower ATE and ARE) compared to state-of-the-art autoregressive and divide-and-conquer baselines, demonstrating stable and high-fidelity generation for long trajectory videos up to 32 seconds in length.

视频生成自回归长视频分治

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。