分析潜空间推理在弱强监督下的表现,发现监督越强越少走捷径但越难保持多样性。
How Do Latent Reasoning Methods Perform Under Weak and Strong Supervision?
- 通过潜空间生成多步推理,避免依赖离散语言符号
- 强监督减少捷径行为但压缩潜表示多样性,弱监督则相反
- 潜空间能存多可能性,但推理过程无结构化搜索
潜空间推理作为一种新型推理范式,通过在连续潜空间中生成推理步骤实现多步推理,突破了离散语言标记的限制。尽管已有大量研究提升其性能,其内部机制仍不明确。本文对不同监督强度下的潜空间推理方法进行系统分析,发现两类共性问题:一是普遍存在捷径行为,高准确率不依赖真正推理;二是虽潜表示可编码多种可能,但推理过程未实现类广度优先搜索的结构化探索,而是表现出隐式剪枝与压缩。研究揭示监督强度存在权衡:强监督抑制捷径但限制潜表示的多样性,弱监督虽增强表示丰富性,却加剧捷径倾向。
原文摘要 · Abstract (English)
Latent reasoning has been recently proposed as a reasoning paradigm and performs multi-step reasoning through generating steps in the latent space instead of the textual space. This paradigm enables reasoning beyond discrete language tokens by performing multi-step computation in continuous latent spaces. Although there have been numerous studies focusing on improving the performance of latent reasoning, its internal mechanisms remain not fully investigated. In this work, we conduct a comprehensive analysis of latent reasoning methods to better understand the role and behavior of latent representation in the process. We identify two key issues across latent reasoning methods with different levels of supervision. First, we observe pervasive shortcut behavior, where they achieve high accuracy without relying on latent reasoning. Second, we examine the hypothesis that latent reasoning supports BFS-like exploration in latent space, and find that while latent representations can encode multiple possibilities, the reasoning process does not faithfully implement structured search, but instead exhibits implicit pruning and compression. Finally, our findings reveal a trade-off associated with supervision strength: stronger supervision mitigates shortcut behavior but restricts the ability of latent representations to maintain diverse hypotheses, whereas weaker supervision allows richer latent representations at the cost of increased shortcut behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。