arXiv:2608.01926cs.AI2026-08

用双曲几何建模视觉世界,让长程目标规划更准。

ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching

论文配图:ProWorld: Progress-Aware Hyperbolic World Models for Long-Horizon Visual Goal Reaching
图 1 · 摘自论文原文
  • 引入进度顺序约束,用双曲空间组织潜在状态演化。
  • 在4个任务上比LeWM提升9.67%成功率,尤其长程任务更优。
  • 适合做长时序视觉导航与目标达成的强化学习研究者。

JEPA风格的视觉世界模型通过预测未来潜在表示来实现视觉目标规划。现有方法通常通过下一步表示预测学习局部转移一致性,但在长程任务中,仅保证局部准确并不足以持续向目标推进。首先,多步回放可能保持局部合理但偏离目标轨迹;其次,局部相似的未来状态可能对应显著不同的长期进展,难以在以局部一致性优化的潜空间中区分。为此,本文提出目标条件下的进度顺序,按状态向目标推进的程度进行相对排序,该顺序具有非对称、粗到精的结构:早期状态保留更广的未来可能性,后期状态聚焦于更具体的目標相关区域。此结构契合双曲几何特性。受此启发,我们提出ProWorld,一种感知进度的双曲视觉世界模型。ProWorld利用目标条件进度顺序组织视觉潜空间动态,通过双曲蕴含学习维持轨迹方向性进展,并借助双曲未来判别缓解局部相似状态间的进展歧义。此外,设计了感知进度的规划目标,联合考虑接近目标程度与中间状态的持续进展来评分候选回放。在四个视觉目标达成任务上的实验表明,ProWorld相比LeWM平均绝对成功率提升9.67%。代码将在论文录用后公开。

原文摘要 · Abstract (English)

JEPA-style visual world models offer an effective paradigm for visual goal planning by predicting future latent representations. Existing methods typically learn local transition consistency through next-step representation prediction. However, in long-horizon tasks, accurate local prediction alone need not ensure sustained progress toward the goal. First, multi-step rollouts can remain locally plausible while drifting away from goal-relevant trajectories. Second, locally similar future states can correspond to substantially different long-term progress, making them difficult to distinguish in a latent space optimized mainly for local consistency. To address these challenges, we introduce goal-conditioned progress order, a relative ordering of states according to how they advance toward a given goal. This order exhibits an asymmetric, coarse-to-fine structure: early states retain broader future possibilities, while later states concentrate on more specific goal-relevant regions. Such a structure is well suited to hyperbolic geometry. Motivated by this observation, we propose ProWorld, a progress-aware hyperbolic visual world model. ProWorld leverages goal-conditioned progress order to organize visual latent-space dynamics, maintains directional progress within trajectories via hyperbolic entailment learning, and mitigates progress ambiguity among locally similar future states via hyperbolic future discrimination. Furthermore, we design a progress-aware planning objective that scores candidate rollouts by jointly considering proximity to the goal and sustained progress across intermediate states. Experiments on four visual goal-reaching tasks demonstrate that ProWorld achieves an average absolute success-rate gain of 9.67 over LeWM. The code will be released after the paper is accepted.

视觉规划双曲空间长程控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。