让3D场景模型预测缺失物体的潜在状态,实现可查询的语义补全。
SR-JEPA: Learning Predictive Latent State in 3D Scenes

- 基于点云设计预测性架构,仅用自包含3D目标训练
- 缺失物体的潜状态达到43.13%语义识别准确率,优于基线22.18点
- 可查询的预测状态融合几何与语义,适合下游3D理解任务
联合嵌入预测架构通过预测缺失观测的潜在表示来学习,但多数掩码式JEPAs仅评估其编码器性能。本文提出SR-JEPA,一种面向场景级点云的点原生JEPA,其预训练的预测路径可在指定位置被查询。评估时,移除某物体所有点后,以无形状的32点查询替代其质心。训练仅使用自包含3D EMA目标:无需重建、语义标签、语言或提升的2D特征。在5,953个独立的ARKitScenes物体上,补全潜状态达43.13%语义身份宏准确率,比最强基线高22.18点;随机化预测路径导致下降9.78点,替换匹配上下文导致下降21.98点。在8,570对Sr3D支持样本上,完整潜状态达41.15 AP;从缺失物体潜状态解码的身份结合锚点身份与几何信息,达39.37 AP,剩余1.78点未解释。结果表明,该模型具备可查询、组合式的3D预测状态,能根据上下文补全实体内容,供下游计算结合度量几何使用。
原文摘要 · Abstract (English)
Joint-embedding predictive architectures learn by predicting latent representations of missing observations, yet many masked JEPAs are evaluated primarily through the encoders they produce. We ask what a trained predictive pathway itself infers when an entire entity is absent from a native 3D scene. We introduce SR-JEPA, a point-native JEPA for scene-scale point clouds whose original frozen predictive pathway can be queried at a supplied location. At evaluation, every point of one object is removed before encoding and replaced by the same shape-free 32-point query at its centroid. Training uses only self-contained 3D EMA targets: no reconstruction, semantic labels, language, or lifted 2D features. On 5,953 held-out ARKitScenes objects, the imputed latent reaches 43.13% semantic-identity macro accuracy, 22.18 points above the strongest floor. Randomizing the prediction path removes 9.78 points, while substituting matched donor context removes 21.98 points. On 8,570 Sr3D support pairs, the full latent reaches 41.15 AP; identity decoded from the missing-object latent, combined with anchor identity and geometry, reaches 39.37 AP, leaving an unresolved 1.78-point residual. These results reveal a queryable, compositional 3D predictive state: the model completes context-dependent entity content, which downstream computation combines with metric geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。