Aether统一建模几何与生成,实现零样本跨域空间推理。
Aether: Geometric-Aware Unified World Modeling
- 联合优化4D动态重建、动作条件预测与目标条件规划
- 无需真实数据训练,合成到真实场景零样本泛化
- 用相机轨迹构建几何感知动作空间,适合机器人视觉任务
将几何重建与生成建模融合,仍是实现类人空间推理的挑战。本文提出Aether,一种统一框架,通过联合优化三项核心能力:(1)4D动态重建,(2)动作条件视频预测,(3)目标条件视觉规划。借助任务交错特征学习,实现重建、预测与规划间的协同知识共享。基于视频生成模型,该框架在未见过真实数据的情况下,仍实现零样本合成到真实场景的泛化。此外,其在动作跟随与重建任务中均表现出零样本泛化能力,归因于内在几何建模。即使无真实数据,重建性能也达到或超过专用模型水平。同时,利用相机轨迹作为几何感知动作空间,提升动作条件预测与视觉规划效果。本工作希望推动物理合理世界建模的新前沿及其应用。
原文摘要 · Abstract (English)
The integration of geometric reconstruction and generative modeling remains a critical challenge in developing AI systems capable of human-like spatial reasoning. This paper proposes Aether, a unified framework that enables geometry-aware reasoning in world models by jointly optimizing three core capabilities: (1) 4D dynamic reconstruction, (2) action-conditioned video prediction, and (3) goal-conditioned visual planning. Through task-interleaved feature learning, Aether achieves synergistic knowledge sharing across reconstruction, prediction, and planning objectives. Building upon video generation models, our framework demonstrates zero-shot synthetic-to-real generalization despite never observing real-world data during training. Furthermore, our approach achieves zero-shot generalization in both action following and reconstruction tasks, thanks to its intrinsic geometric modeling. Notably, even without real-world data, its reconstruction performance is comparable with or even better than that of domain-specific models. Additionally, Aether employs camera trajectories as geometry-informed action spaces, enabling effective action-conditioned prediction and visual planning. We hope our work inspires the community to explore new frontiers in physically-reasonable world modeling and its applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。