arXiv:2603.02263cs.CVcs.AI2026-03

多个视角的智能体通过预测学习自发形成可线性转换的隐空间,实现零参数迁移。

Social-JEPA: Emergent Geometric Isomorphism

  • 各视角智能体独立训练,无参数共享
  • 隐空间间存在近似线性等距关系,支持跨视角映射
  • 可实现零梯度迁移,显著降低计算开销

世界模型将丰富的感官流压缩为紧凑的隐码以预测未来观测。我们让多个独立智能体从同一环境的不同视角获取这些模型,且不共享参数或进行协调。训练后,它们的内部表示展现出显著的涌现特性:两个隐空间之间存在近似线性等距关系,可在不同视角间实现透明转换。该几何一致性在大视角偏移和原始像素重叠极少的情况下依然保持。利用学习到的对齐,一个智能体上训练的分类器可直接迁移到另一智能体,无需额外梯度更新;类似蒸馏的迁移加速后续学习并大幅减少总计算量。研究揭示,预测学习目标对表示几何施加强正则性,为去中心化视觉系统间的轻量级互操作提供路径。代码已公开于 https://anonymous.4open.science/r/Social-JEPA-5C57。

原文摘要 · Abstract (English)

World models compress rich sensory streams into compact latent codes that anticipate future observations. We let separate agents acquire such models from distinct viewpoints of the same environment without any parameter sharing or coordination. After training, their internal representations exhibit a striking emergent property: the two latent spaces are related by an approximate linear isometry, enabling transparent translation between them. This geometric consensus survives large viewpoint shifts and scant overlap in raw pixels. Leveraging the learned alignment, a classifier trained on one agent can be ported to the other with no additional gradient steps, while distillation-like migration accelerates later learning and markedly reduces total compute. The findings reveal that predictive learning objectives impose strong regularities on representation geometry, suggesting a lightweight path to interoperability among decentralized vision systems. The code is available at https://anonymous.4open.science/r/Social-JEPA-5C57.

自监督隐空间对齐跨视角迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。