arXiv:2602.02393cs.CVcs.AI2026-02被引 34

让世界模型在真实环境里记住上千帧画面,不依赖精确姿态数据。

Infinite-World: Scaling Interactive World Models to 1000-Frame Horizons via Pose-Free Hierarchical Memory

  • 用分层无姿态记忆压缩器,自动保存长期视觉记忆
  • 在真实视频上实现1000+帧的视觉连贯性,质量显著提升
  • 适合做长时交互的机器人或虚拟世界系统开发

我们提出Infinite-World,一种能在复杂真实环境中小时级保持视觉记忆的交互式世界模型。现有模型虽可在合成数据上高效训练,但因真实视频中姿态估计噪声大、视角复现少,难以有效训练。为此,我们设计分层无姿态记忆压缩器(HPMC),递归将历史隐状态压缩为固定预算表示;联合优化生成主干与压缩器,使模型能以可控开销自主锚定远期生成,无需显式几何先验。其次,提出不确定性感知动作标注模块,将连续运动离散化为三态逻辑,最大化原始视频利用率,同时避免噪声轨迹污染确定性动作空间。此外,基于初步实验洞察,采用30分钟紧凑数据集进行密集回访微调,高效激活模型长程环闭合能力。大量实验(含客观指标与用户评估)表明,Infinite-World在视觉质量、动作可控性和空间一致性上均表现卓越。

原文摘要 · Abstract (English)

We propose Infinite-World, a robust interactive world model capable of maintaining coherent visual memory over 1000+ frames in complex real-world environments. While existing world models can be efficiently optimized on synthetic data with perfect ground-truth, they lack an effective training paradigm for real-world videos due to noisy pose estimations and the scarcity of viewpoint revisits. To bridge this gap, we first introduce a Hierarchical Pose-free Memory Compressor (HPMC) that recursively distills historical latents into a fixed-budget representation. By jointly optimizing the compressor with the generative backbone, HPMC enables the model to autonomously anchor generations in the distant past with bounded computational cost, eliminating the need for explicit geometric priors. Second, we propose an Uncertainty-aware Action Labeling module that discretizes continuous motion into a tri-state logic. This strategy maximizes the utilization of raw video data while shielding the deterministic action space from being corrupted by noisy trajectories, ensuring robust action-response learning. Furthermore, guided by insights from a pilot toy study, we employ a Revisit-Dense Finetuning Strategy using a compact, 30-minute dataset to efficiently activate the model's long-range loop-closure capabilities. Extensive experiments, including objective metrics and user studies, demonstrate that Infinite-World achieves superior performance in visual quality, action controllability, and spatial consistency.

世界模型长时记忆真实视频交互系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。