测试大模型能否预判自身回答特性,揭示其自我认知缺陷
The Self-Execution Benchmark: Measuring LLMs' Attempts to Overcome Their Lack of Self-Execution
- 设计自执行基准,评估模型预判自身输出能力
- 多数模型在预判自身难答、拒答等行为上表现不佳
- 模型越大越强不等于自知能力越强,适合研究模型可解释性的人
大型语言模型(LLMs)通常在知识或推理任务上进行评估。本文探索一种新型评测方式:模型能否预测自身回答的某些特征。由于LLMs缺乏自我执行能力,我们提出自执行基准(Self-Execution Benchmark),衡量模型对自身输出属性的预见能力,例如问题是否对其困难、是否会拒绝回答,或可能产生何种关联。实验表明,模型在该基准上整体表现较差,且模型规模或能力提升并不总带来性能改善。结果表明,LLMs在表征和推理自身行为方面存在根本性局限。
原文摘要 · Abstract (English)
Large language models (LLMs) are commonly evaluated on tasks that test their knowledge or reasoning abilities. In this paper, we explore a different type of evaluation: whether an LLM can predict aspects of its own responses. Since LLMs lack the ability to execute themselves, we introduce the Self-Execution Benchmark, which measures a model's ability to anticipate properties of its output, such as whether a question will be difficult for it, whether it will refuse to answer, or what kinds of associations it is likely to produce. Our experiments show that models generally perform poorly on this benchmark, and that increased model size or capability does not consistently lead to better performance. These results suggest a fundamental limitation in how LLMs represent and reason about their own behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。