arXiv:2608.30980cs.CLcs.AI2026-09

提升大模型自我认知能力,让其能准确判断自身行为变化。

Evaluating and Improving LLM Self-Modeling

论文配图:Evaluating and Improving LLM Self-Modeling
图 1 · 摘自论文原文
  • 构建多样自模拟能力评测基准,测试模型对自身行为的判断。
  • 强化学习结合合成数据,显著提升三类开源模型的自模拟能力。
  • 改进后的模型仍缺乏深层内省,非源于内部决策信息访问。

我们研究自模型化:大语言模型对其自身行为进行问答的能力。重点关注可验证的行为问题,例如提示词修改是否会影响模型最终答案。为衡量该能力,我们引入一个涵盖多种类型自模型化问题的基准测试。当前模型展现出非平凡但有限的自模型化技能,并在关于自身行为的简单反事实问题上系统性出错。为提升自模型化能力,我们开发了一种可扩展的合成数据生成流程,用于生成自模型化训练数据,并证明强化学习可提升三个开源模型家族的整体自模型化能力,且在未见任务上具有部分迁移效果。然而,这些提升并未体现为一致的内省:自模型化能力的改善并非源于对模型内部决策过程的特权访问。

原文摘要 · Abstract (English)

We study self-modeling: an LLM's ability to answer questions about its own behavior. We focus on verifiable behavioral questions, such as whether a prompt edit would change the model's final answer. To measure this capability, we introduce a benchmark that tests diverse types of self-modeling questions. Current models show non-trivial but limited self-modeling skill, and make systematic mistakes on simple counterfactual questions about their own behavior. To improve self-modeling skill, we develop a scalable synthetic-data pipeline that produces self-modeling training data, and show that reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks. These gains, however, do not seem to constitute introspection consistently: improved self-modeling may not arise from privileged access to the model's internal decision process.

自模型化LLM评估强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。