arXiv:2504.10615cs.CL2025-04被引 4

测试大模型在隐空间中的推理能力,发现其能无声思考但存在潜在风险。

Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models

  • 用非英语初始词选择方式强制模型内部推理
  • GPT-4.5准确率达74.7%,显著高于其他模型
  • 揭示模型隐性推理机制,警示安全风险

大型语言模型(LLMs)既能通过生成显式思维链进行外部推理,也能在隐空间内进行内部推理。尽管通过扩展测试时计算量已取得显著进展,但对模型内部推理能力的理解与量化仍至关重要。本研究构建了一个包含4,000个样本的基准,用于量化不同领域的模型内部推理能力。方法是要求模型不通过描述性文本,而是通过选择初始响应词的语言(不同于提示语言)来给出正确答案,这不仅要求模型超越上下文窗口,还需克服默认使用相同语言的倾向,增加认知负担。评估了18个模型,结果差异显著,GPT-4.5准确率最高(74.7%),优于Grok-2(67.2%)和Llama 3.1 405B(65.6%)。控制实验与难度扩展分析表明,尽管模型存在内部推理,但在某些条件下仍可能利用启发式策略,需进一步研究。实验表明,模型可通过隐空间计算“思考”,揭示了需深入理解的内部推理策略,尤其涉及隐蔽规划、目标追求或欺骗等安全问题。

原文摘要 · Abstract (English)

Large language models (LLMs) can perform reasoning computations both internally within their latent space and externally by generating explicit token sequences like chains of thought. Significant progress in enhancing reasoning abilities has been made by scaling test-time compute. However, understanding and quantifying model-internal reasoning abilities - the inferential "leaps" models make between individual token predictions - remains crucial. This study introduces a benchmark (n = 4,000 items) designed to quantify model-internal reasoning in different domains. We achieve this by having LLMs indicate the correct solution to reasoning problems not through descriptive text, but by selecting a specific language of their initial response token that is different from English, the benchmark language. This not only requires models to reason beyond their context window, but also to overrise their default tendency to respond in the same language as the prompt, thereby posing an additional cognitive strain. We evaluate a set of 18 LLMs, showing significant performance variations, with GPT-4.5 achieving the highest accuracy (74.7%), outperforming models like Grok-2 (67.2%), and Llama 3.1 405B (65.6%). Control experiments and difficulty scaling analyses suggest that while LLMs engage in internal reasoning, we cannot rule out heuristic exploitations under certain conditions, marking an area for future investigation. Our experiments demonstrate that LLMs can "think" via latent-space computations, revealing model-internal inference strategies that need further understanding, especially regarding safety-related concerns such as covert planning, goal-seeking, or deception emerging without explicit token traces.

模型推理隐空间安全风险评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。