Astra能精准预测多种任务的长期未来,支持实时交互控制。
Astra: General Interactive World Model with Autoregressive Denoising
- 采用自回归去噪架构与时间因果注意力,实现流式输出与长程一致性。
- 在多个数据集上优于现有模型,在动作对齐和长期预测上提升显著。
- 适合自动驾驶、机器人操作等需要精确动作交互的通用场景。
扩散变压器的进展使文本或图像生成高质量视频成为可能,但能够从历史观测和动作中预测长期未来的通用世界模型仍研究不足,尤其在多样化任务和动作形式下。为此,我们提出Astra,一种交互式通用世界模型,可在多种场景(如自动驾驶、机器人抓取)中生成包含精确动作交互(如相机运动、机械臂动作)的真实未来视频。采用自回归去噪架构并引入时间因果注意力以聚合历史信息并支持流式输出。通过噪声增强的历史记忆避免过度依赖过去帧,平衡响应性与时间连贯性。为实现精准动作控制,设计动作感知适配器,直接将动作信号注入去噪过程;进一步构建动作专家混合模型,动态路由异构动作模态,提升在探索、操作、摄像控制等多样任务中的泛化能力。Astra在交互性、一致性与通用性方面表现优异,实验表明其在保真度、长程预测与动作对齐上均超越现有最优世界模型。
原文摘要 · Abstract (English)
Recent advances in diffusion transformers have empowered video generation models to generate high-quality video clips from texts or images. However, world models with the ability to predict long-horizon futures from past observations and actions remain underexplored, especially for general-purpose scenarios and various forms of actions. To bridge this gap, we introduce Astra, an interactive general world model that generates real-world futures for diverse scenarios (e.g., autonomous driving, robot grasping) with precise action interactions (e.g., camera motion, robot action). We propose an autoregressive denoising architecture and use temporal causal attention to aggregate past observations and support streaming outputs. We use a noise-augmented history memory to avoid over-reliance on past frames to balance responsiveness with temporal coherence. For precise action control, we introduce an action-aware adapter that directly injects action signals into the denoising process. We further develop a mixture of action experts that dynamically route heterogeneous action modalities, enhancing versatility across diverse real-world tasks such as exploration, manipulation, and camera control. Astra achieves interactive, consistent, and general long-term video prediction and supports various forms of interactions. Experiments across multiple datasets demonstrate the improvements of Astra in fidelity, long-range prediction, and action alignment over existing state-of-the-art world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。