arXiv:2601.12269cs.CLcs.AI2026-01被引 2

用模拟退火提升语言模型的理论心智推理能力

Simulated Annealing Enhances Theory-of-Mind Reasoning in Autoregressive Language Models

  • 通过马尔可夫链蒙特卡洛采样优化序列概率分布
  • 渐进降温的模拟退火使理论心智任务准确率显著提升
  • 无需微调即可激活模型隐含推理能力,适合追求高效推理的研究者

自回归语言模型是下一个词预测器,常被批评为仅优化表面合理性(局部连贯性),而非保持正确的潜在状态表征(全局连贯性)。由于理论心智(ToM)任务高度依赖对自身与他人潜在心理状态的推理,此类模型通常被认为无法胜任。尽管后训练方法可提升其表现,我们证明:无需任何权重更新或验证,仅通过基础模型即可恢复强理论心智能力。本方法基于近期幂采样技术(Karan & Du, 2025),利用马尔可夫链蒙特卡洛(MCMC)从自回归语言模型的尖锐化序列级(非词元级)概率分布中采样。进一步发现,引入模拟退火——将温度从高到低逐步降低——相较于固定温度的幂采样,显著提升了理论心智表现。结果表明,基于采样的优化为在不重新训练的前提下挖掘语言模型隐含能力提供了有效路径。

原文摘要 · Abstract (English)

Autoregressive language models are next-token predictors and have been criticized for only optimizing surface plausibility (i.e., local coherence) rather than maintaining correct latent-state representations (i.e., global coherence). Because Theory of Mind (ToM) tasks crucially depend on reasoning about latent mental states of oneself and others, such models are therefore often thought to fail at ToM. While post-training methods can improve ToM performance, we show that strong ToM capability can be recovered directly from the base model without any additional weight updates or verifications. Our approach builds on recent power-sampling methods (Karan & Du, 2025) that use Markov chain Monte Carlo (MCMC) to sample from sharpened sequence-level (rather than token-level) probability distributions of autoregressive language models. We further find that incorporating annealing, where the tempered distribution is gradually shifted from high to low temperature, substantially improves ToM performance over fixed-temperature power sampling. Together, these results suggest that sampling-based optimization provides a powerful way to extract latent capabilities from language models without retraining.

理论心智采样优化模拟退火

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。