将内在能量与联合嵌入结合,构建具方向性的状态空间距离
Intrinsic-Energy Joint Embedding Predictive Architectures Induce Quasimetric Spaces
- 用最小作用量原理定义内在能量,生成有向距离
- 内在能量天然满足非对称性,适配单向可达场景
- 为强化学习中的目标导向控制提供理论支持
联合嵌入预测架构(JEPAs)通过从上下文嵌入预测目标嵌入,在潜在空间中诱导出标量兼容性能量。而拟度量强化学习(QRL)则通过有向距离值(到达目标的代价)来实现目标条件控制,适用于非对称动态系统。本文将两者结合,聚焦于一类合理的JEPA能量函数:内在(最小作用量)能量,定义为两点间可接受轨迹上局部努力累积的下确界。在温和的闭包与可加性假设下,任意内在能量均为拟度量。在目标可达控制中,最优代价到达函数恰好具有此内在形式;反之,训练以建模内在能量的JEPA,其输出属于QRL所追求的拟度量值类别。此外,我们指出对称有限能量在单向可达性中存在结构不匹配,因此当方向性重要时,应采用非对称(拟度量)能量。
原文摘要 · Abstract (English)
Joint-Embedding Predictive Architectures (JEPAs) aim to learn representations by predicting target embeddings from context embeddings, inducing a scalar compatibility energy in a latent space. In contrast, Quasimetric Reinforcement Learning (QRL) studies goal-conditioned control through directed distance values (cost-to-go) that support reaching goals under asymmetric dynamics. In this short article, we connect these viewpoints by restricting attention to a principled class of JEPA energy functions : intrinsic (least-action) energies, defined as infima of accumulated local effort over admissible trajectories between two states. Under mild closure and additivity assumptions, any intrinsic energy is a quasimetric. In goal-reaching control, optimal cost-to-go functions admit exactly this intrinsic form ; inversely, JEPAs trained to model intrinsic energies lie in the quasimetric value class targeted by QRL. Moreover, we observe why symmetric finite energies are structurally mismatched with one-way reachability, motivating asymmetric (quasimetric) energies when directionality matters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。