arXiv:2602.01148cs.AIcs.IT2026-02被引 3

揭示潜空间推理的探索与执行权衡机制,提出动态调节决策确定性的新方法。

Capabilities and Fundamental Limits of Latent Chain-of-Thought

  • 用符号索引量化决策确定性,解释模型在探索与执行间的性能波动
  • 在ProsQA上达97.0%准确率,在GSM8K上仅34.1%,体现探索-执行权衡
  • 适合研究高效推理架构、持续表示模型的学者参考

潜空间链式思维(Latent CoT)模型通过连续表示实现高效推理,但表现出令人困惑的性能不一致:在探索任务中表现优异(ProsQA: 97.0%),在计算任务中却表现不佳(GSM8K: 34.1%)。我们揭示这种权衡由决策确定性决定。贡献有三:(1) 理论上刻画了根本性的探索-执行权衡,证明高确定性利于精确执行但抑制探索,低确定性促进搜索但导致误差累积;(2) 引入符号索引——量化决策承诺的核心机制,建立其与执行稳定性及探索能力的因果关系;(3) 证明课程学习理论上必要,因直接训练会因分布不匹配而失败。该框架将设计范式从二元架构选择转向根据任务需求动态调节决策确定性的自适应系统。

原文摘要 · Abstract (English)

Latent Chain-of-Thought (Latent CoT) models promise efficient reasoning via continuous representations, yet exhibit puzzling performance inconsistencies: excelling at exploration (ProsQA: 97.0%) but failing at computation (GSM8K: 34.1%). We reveal that this trade-off is governed by decisional certainty. Our contributions are threefold: (1) We theoretically characterize the fundamental Exploration-Execution Trade-off, proving that high certainty enables precise execution but inhibits exploration, while low certainty facilitates search but causes error accumulation. (2) We introduce the Symbolic Index--quantifying decisional commitment--as the core mechanism governing this trade-off and establish its causal relationship with both execution stability and exploration capability. (3) We prove that curriculum learning is theoretically necessary, as direct training provably fails due to distributional mismatch. Our framework shifts the design paradigm from binary architectural choices toward adaptive systems that dynamically regulate decisional certainty based on task demands.

潜空间推理链式思维决策确定性自适应系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。