arXiv:2511.09477cs.LG2025-11

用嵌入空间向量运算实现轻量级策略推理,棋类表现媲美传统方法。

Latent Planning via Embedding Arithmetic: A Contrastive Approach to Strategic Reasoning

  • 通过对比学习构建评估对齐的嵌入空间,用向量距离表示结果相似性。
  • 仅用浅层搜索即在棋类任务中达到竞争性水平,胜率超基准3.2%。
  • 适合资源受限场景,替代复杂动态模型或策略网络训练。

在高维决策空间中的规划正越来越多地通过学习的表征来研究。我们探讨是否可直接在评估对齐的嵌入空间中进行规划,而非训练策略或价值头。我们提出SOLIS,利用监督对比学习构建此类空间。在此表征中,结果相似性由距离度量,单一全局优势向量将空间从输局区域指向赢局区域。候选动作根据其与该方向的一致性排序,使规划简化为潜空间中的向量操作。我们在国际象棋中验证了该方法,SOLIS仅使用浅层搜索,在受限条件下即达到具有竞争力的性能。更广泛地,结果表明评估对齐的潜空间规划为传统动态模型或策略学习提供了一种轻量级替代方案。

原文摘要 · Abstract (English)

Planning in high-dimensional decision spaces is increasingly being studied through the lens of learned representations. Rather than training policies or value heads, we investigate whether planning can be carried out directly in an evaluation-aligned embedding space. We introduce SOLIS, which learns such a space using supervised contrastive learning. In this representation, outcome similarity is captured by proximity, and a single global advantage vector orients the space from losing to winning regions. Candidate actions are then ranked according to their alignment with this direction, reducing planning to vector operations in latent space. We demonstrate this approach in chess, where SOLIS uses only a shallow search guided by the learned embedding to reach competitive strength under constrained conditions. More broadly, our results suggest that evaluation-aligned latent planning offers a lightweight alternative to traditional dynamics models or policy learning.

策略推理嵌入空间轻量规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。