arXiv:2608.29434cs.LGcs.AI2026-08

将JEPA世界模型从图像扩展到点云,实现几何观测下的稳定规划。

Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations

论文配图:Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations
图 1 · 摘自论文原文
  • 基于点云设计三类JEPA模型,保持潜在空间规划能力。
  • 在30%点移动场景中仍能稳定规划,最优模型成功率超图像基线。
  • 直接用目标点云构造目标潜在表示,无需额外观察,适配机器人控制。

JEPA世界模型使潜在空间规划成为控制的可行路径,但其几乎全基于图像构建。点云具有稀疏、无序、自遮挡等特性,且仅0.3%-15%的场景点移动时,潜在预测的慢特征最优性与3D自监督的几何捷径叠加,使潜在预测是否有效尚不明确。本文将三种经典JEPA架构(冻结编码器、分布先验、动作敏感)迁移至点云,重新评估稳定世界模型基准,仅改变观测模态。三类模型均未出现崩溃:分布先验模型在所有基准上与重评估的图像对应模型统计等效;动作敏感模型在几何变化最剧烈的场景中表现最佳。探查表明:物体位置可近乎线性解码,注意力集中于少数移动点。规划对训练中未见的强数据缺失具有鲁棒性,但范围噪声会破坏最稀疏场景。几何结构最终使指定3D目标成为自然的目标接口:通过目标点云与当前状态的潜在表示构造目标潜在向量,成功率不变,无需额外目标观测。

原文摘要 · Abstract (English)

JEPA world models make latent-space planning a practical route to control, but they are built almost exclusively on images. Whether latent prediction survives geometric observations is unclear: point clouds are sparse, unordered, and self-occluded, and with 0.3-15% of scene points moving, the slow-feature optimum of latent prediction compounds with the geometric shortcut of 3D self-supervision. We lift three canonical JEPA designs to point clouds, frozen-encoder, distribution-prior, and action-sensitive, and re-sense the stable-worldmodel benchmark so that only the observation differs from the image baselines. All three plan without collapse: the distribution-prior model is statistically equivalent to its re-evaluated image counterpart on every benchmark, and the action-sensitive model attains the strongest result in our controlled comparison where the most geometry moves. Probing explains why: object positions are almost perfectly linearly decodable and attention falls on the few moving points. Planning withstands heavy dropout never seen in training, though range noise defeats the thinnest scene. Geometry finally makes a commanded 3D target a natural goal interface: we construct the goal latent from the target and the current latent, at no cost in success rate, without a goal observation.

世界模型点云潜空间规划3D控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。