arXiv:2602.22960cs.CV2026-02International Conf…被引 4

统一建模相机控制与长期记忆,实现高保真视频生成

UCM: Unified Modeling of Camera Control and Memory with Time-aware Positional Encoding Warping for World Models

  • 用时间感知位置编码扭曲统一建模相机控制与记忆
  • 在50万+单目视频上训练,长时一致性显著提升
  • 适合需要精准相机控制的交互式环境模拟场景

基于视频生成的世界模型在模拟交互环境方面展现出巨大潜力,但在场景重访时保持长期内容一致性和实现用户指定输入下的精确相机控制方面仍面临挑战。现有方法依赖显式三维重建,灵活性受限且难以保留细粒度结构;另一类方法直接依赖先前生成帧,缺乏显式空间对应关系,限制了可控性与一致性。为此,我们提出UCM框架,通过时间感知位置编码扭曲机制,统一建模长期记忆与精确相机控制。为降低计算开销,设计高效双流扩散变压器以实现高保真生成。此外,引入可扩展的数据构建策略,利用点云渲染模拟场景重访,支持在超过50万条单目视频上训练。在真实世界与合成基准上的大量实验表明,UCM在长时场景一致性上显著优于现有方法,同时在高保真视频生成中实现了精确相机控制。

原文摘要 · Abstract (English)

World models based on video generation demonstrate remarkable potential for simulating interactive environments yet suffer from persistent difficulties in two key areas: maintaining long-term content consistency when scenes are revisited and enabling precise camera control from user-specified inputs. Existing methods based on explicit 3D reconstruction often compromise flexibility in unbounded scenarios and struggle to preserve fine-grained structures. Alternative methods rely directly on previously generated frames without establishing explicit spatial correspondence, thereby limiting controllability and consistency. To address these limitations, we present UCM, a novel framework for unified modeling of long-term memory and precise camera control via a time-aware positional encoding warping mechanism. To reduce computational overhead, we design an efficient dual-stream diffusion transformer for high-fidelity generation. Moreover, we introduce a scalable data curation strategy that utilizes point-cloud-based rendering to simulate scene revisiting, enabling training on over 500K monocular videos. Extensive experiments on real-world and synthetic benchmarks demonstrate that UCM significantly outperforms state-of-the-art methods on long-term scene consistency, while achieving precise camera controllability in high-fidelity video generation.

世界模型视频生成相机控制长时一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。