用双曲几何提升视觉预测模型的长期规划能力。
GeoWorld: Geometric World Models
- 将隐变量映射到双曲空间,保留状态间的几何与层次结构。
- 在3步和4步规划中分别提升约3%和2%的成功率。
- 适合需要长序列推理的强化学习与视觉规划任务。
基于能量的预测性世界模型通过推理潜在能量景观而非生成像素,实现了多步视觉规划。然而现有方法存在两大挑战:(i) 隐变量通常在欧几里得空间中学习,忽略了状态间的内在几何与层次结构;(ii) 长时程预测性能迅速退化。为此,我们提出GeoWorld,一种通过双曲JEPA将隐变量从欧几里得空间映射至双曲流形以保持几何结构与层次关系的世界模型。我们进一步引入几何强化学习进行能量优化,实现双曲潜在空间中的稳定多步规划。在CrossTask和COIN上的大量实验表明,相较于最先进方法V-JEPA 2,GeoWorld在3步规划中提升约3%成功率,在4步规划中提升约2%。项目网站:https://steve-zeyu-zhang.github.io/GeoWorld。
原文摘要 · Abstract (English)
Energy-based predictive world models provide a powerful approach for multi-step visual planning by reasoning over latent energy landscapes rather than generating pixels. However, existing approaches face two major challenges: (i) their latent representations are typically learned in Euclidean space, neglecting the underlying geometric and hierarchical structure among states, and (ii) they struggle with long-horizon prediction, which leads to rapid degradation across extended rollouts. To address these challenges, we introduce GeoWorld, a geometric world model that preserves geometric structure and hierarchical relations through a Hyperbolic JEPA, which maps latent representations from Euclidean space onto hyperbolic manifolds. We further introduce Geometric Reinforcement Learning for energy-based optimization, enabling stable multi-step planning in hyperbolic latent space. Extensive experiments on CrossTask and COIN demonstrate around 3% SR improvement in 3-step planning and 2% SR improvement in 4-step planning compared to the state-of-the-art V-JEPA 2. Project website: https://steve-zeyu-zhang.github.io/GeoWorld.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。