arXiv:2606.27964cs.CV2026-06被引 1

实现人姿与镜头运动组合控制的快速视频生成,支持长时序稳定输出。

Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control

论文配图:Directing the World: Fast Autoregressive Video Generation with Compositional Human-Camera Control
图 1 · 摘自论文原文
  • 解耦控制学习,保持统一自回归视频先验,提升生成稳定性。
  • 支持多人群体动作与视角变化组合,生成视频时长超10秒仍清晰可控。
  • 构建大规模同步标注数据集,适用于人像与镜头协同控制任务。

构建交互式世界模型需要在长时序下生成逼真视频并保持可控动态。自回归视频生成具备可扩展性,但长期推演时易出现误差累积和时间退化问题。当涉及人体动作与摄像机轨迹等异构控制时,现有方法常因干扰导致预训练视频先验不稳定,且难以兼顾可控性与画质。本文提出「Directing the World」框架,实现基于人体动作与镜头轨迹组合控制的快速自回归视频生成。核心思想是解耦控制学习,同时保留统一的自回归视频先验。引入快慢记忆训练策略以稳定长时序生成并加速收敛。针对人体动作控制,设计时间引导的动态投影机制与优化的运动CFG策略,实现时序平滑的动作对齐,不牺牲视觉质量,并支持多人控制。在学习鲁棒动作先验后,引入第二阶段镜头轨迹控制模块,将人体动态与视角变化融合,实现连贯的世界探索。进一步构建大规模同步数据集,包含视频、文本、人体动作与镜头轨迹标注,按动作中心与镜头中心组织,支持解耦训练。大量实验表明,该方法在超过10秒的长时序生成中仍保持稳定可控与高质量输出。

原文摘要 · Abstract (English)

Building interactive world models requires generating realistic videos while maintaining controllable dynamics over long horizons. Autoregressive video generation offers a scalable foundation, but suffers from error accumulation and temporal degradation during extended rollouts. This issue is further amplified under heterogeneous controls such as human motion and camera trajectories, which may interfere and destabilize a pretrained video prior, while existing methods often trade off controllability and visual quality. We propose "Directing the World", a fast autoregressive framework for controllable world-model video generation with compositional human-motion and camera-trajectory control. Our key idea is to decouple control learning while preserving a unified autoregressive video prior. We introduce a Fast-Slow Memory training strategy to stabilize long-horizon rollout learning and improve convergence. For human motion control, we design a t-guided Dynamic Projection mechanism and a refined Motion-CFG strategy, enabling temporally smooth and accurate motion alignment without degrading visual fidelity, and supporting multi-person control.After learning a robust motion prior, we introduce a second-stage camera-trajectory control module to compose human dynamics with viewpoint changes for coherent world exploration. We further construct a large-scale dataset with synchronized video, text, human-motion, and camera-trajectory annotations, organized into motion-centric and camera-centric subsets for decoupled training. Extensive experiments show stable long-horizon generation with precise controllability and high visual quality. See more at https://whydahuzi.github.io/Directing-the-World.github.io/.

视频生成自回归动作控制镜头控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。