arXiv:2603.20327cs.LGcs.AI2026-03

发现视频模型隐空间中隐藏的离散符号与物理结构

Probing the Latent World: Emergent Discrete Symbols and Physical Structure in Latent Representations

  • 用无监督量化方法将连续隐向量转为离散符号序列
  • 在Kinetics-mini上验证了三种物理维度的显著差异
  • 适合研究模型可解释性与符号化世界模型的学者

基于联合嵌入预测架构(JEPA)训练的视频世界模型通过预测隐空间中的掩码区域获得丰富的时空表征,而非重建像素。这消除了生成模型的视觉验证路径,造成结构可解释性缺口:编码器学习到的物理结构无法直接观测。现有探测方法要么在连续空间操作,缺乏结构中间层;要么附加生成组件,导致行为归因混淆。本文提出被动量化探测框架AI母语(AIM),无需任务监督、不修改编码器,将V-JEPA 2的连续隐向量转换为离散符号序列。因编码器完全冻结,任何符号结构均源自预训练表示。在Kinetics-mini上通过类别对比实验,检验抓取角度、物体几何与运动时间结构三个物理维度。结果显示,不同类别间符号分布差异显著(卡方检验p < 10^{-4};互信息0.036–0.117比特,标准化互信息1.2–3.9%最大值;杰普森散度最高0.342;代码本活跃率62.5%)。结果表明,V-JEPA 2隐空间高度紧凑:多样动作类别共享共同表征核心,语义差异以分布梯度变化呈现而非分类边界。该工作确立了面向动作条件符号化世界模型四阶段路线图的第一阶段,证明结构化符号流形是冻结JEPA隐空间可发现的属性。

原文摘要 · Abstract (English)

Video world models trained with Joint Embedding Predictive Architectures (JEPA) acquire rich spatiotemporal representations by predicting masked regions in latent space rather than reconstructing pixels. This removes the visual verification pathway of generative models, creating a structural interpretability gap: the encoder has learned physical structure inaccessible in any inspectable form. Existing probing methods either operate in continuous space without a structured intermediate layer, or attach generative components whose parameters confound attribution of behavior to the encoder. We propose the AI Mother Tongue (AIM) framework as a passive quantization probe: a lightweight, vocabulary-free probe that converts V-JEPA 2 continuous latent vectors into discrete symbol sequences without task-specific supervision or modifying the encoder. Because the encoder is kept completely frozen, any symbolic structure in the AIM codebook is attributable entirely to V-JEPA 2 pre-trained representations -- not to the probe. We evaluate through category-contrast experiments on Kinetics-mini along three physical dimensions: grasp angle, object geometry, and motion temporal structure. AIM symbol distributions differ significantly across all three experiments (chi^2 p < 10^{-4}; MI 0.036--0.117 bits, NMI 1.2--3.9% of the 3-bit maximum; JSD up to 0.342; codebook active ratio 62.5%). The experiments reveal that V-JEPA 2 latent space is markedly compact: diverse action categories share a common representational core, with semantic differences encoded as graded distributional variations rather than categorical boundaries. These results establish Stage 1 of a four-stage roadmap toward an action-conditioned symbolic world model, demonstrating that structured symbolic manifolds are discoverable properties of frozen JEPA latent spaces.

隐空间探查符号化视频建模可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。