评测大模型在代理任务中自我推理的能力,发现仅顶级模型具备此能力。
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
- 设计多场景评估任务,测试模型自我修改、知识获取和隐蔽推理能力。
- 仅最先进模型表现出自我推理能力,且结果高度依赖上下文。
- 评估工具开源,可用于监测未来模型的自我推理进步程度。
我们提出一套任务,用于评估大型语言模型(LLM)代理的工具性自我推理能力。该能力可提升模型适应性并实现自我修改,但也可能引发欺骗性对齐等风险。以往研究仅限于非代理场景或有限领域。本文在广泛场景下评估代理任务中的工具性自我推理,包括自我修改、知识获取和不透明推理。我们评估了使用前沿LLM构建的代理,涵盖商用与开源系统。结果发现,只有最强大的前沿模型才表现出该能力,且其表现高度依赖上下文。所有模型均未通过最困难的评估版本,因此该评估可用来衡量未来模型在工具性自我推理上的进步。评估代码已开源:https://github.com/kaifronsdal/Self-Reasoning-Evals。
原文摘要 · Abstract (English)
We propose a suite of tasks to evaluate the instrumental self-reasoning ability of large language model (LLM) agents. Instrumental self-reasoning ability could improve adaptability and enable self-modification, but it could also pose significant risks, such as enabling deceptive alignment. Prior work has only evaluated self-reasoning in non-agentic settings or in limited domains. In this paper, we propose evaluations for instrumental self-reasoning ability in agentic tasks in a wide range of scenarios, including self-modification, knowledge seeking, and opaque self-reasoning. We evaluate agents built using state-of-the-art LLMs, including commercial and open source systems. We find that instrumental self-reasoning ability emerges only in the most capable frontier models and that it is highly context-dependent. No model passes the the most difficult versions of our evaluations, hence our evaluation can be used to measure increases in instrumental self-reasoning ability in future models. We open-source our evaluations at https://github.com/kaifronsdal/Self-Reasoning-Evals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。