大模型常高估自己能力,但部分能从失败中学习改进决策。
Do Large Language Models Know What They Are Capable Of?
- 测试模型是否能预判任务成功率及多步任务中的表现变化。
- 多数模型过自信,但具备优于随机的判断能力,新旧模型无明显差异。
- 部分模型通过失败经验降低过度自信,适合研究自主决策与安全对齐。
我们研究大语言模型(LLMs)能否预测自身在特定任务中的成功概率,以及在多步骤任务中进展时预测能力是否提升。还考察了在失败代价高昂的情境下,模型能否通过上下文经验学习,做出更优的任务取舍决策。所有测试的模型均表现出过度自信,但多数具备优于随机的辨别能力。尽管较新或较大的模型通常未展现出更强的辨别力,不过Claude系列显示出这一趋势。在多步代理任务中,前沿模型的过度自信随任务推进而加剧,推理型模型的表现与非推理型模型相当或更差。通过上下文失败经历,部分模型能减少过度自信,显著改善决策,而另一些则不能。有趣的是,所有模型的决策在给定其成功概率估计下大致理性,但因过于乐观导致实际决策效果不佳。结果表明,当前大模型代理受限于对其自身能力的认知不足。我们讨论了模型对自身能力认知对人工智能滥用和对齐风险的影响。
原文摘要 · Abstract (English)
We investigate whether large language models (LLMs) can predict whether they will succeed on a given task and whether their predictions improve as they progress through multi-step tasks. We also investigate whether LLMs can learn from in-context experiences to make better decisions about whether to pursue a task in scenarios where failure is costly. All LLMs we tested are overconfident, but most predict their success with better-than-random discriminatory power. We find that newer and larger LLMs generally do not have greater discriminatory power, though Claude models do show such a trend. On multi-step agentic tasks, the overconfidence of several frontier LLMs worsens as they progress through the tasks, and reasoning LLMs perform comparably to or worse than non-reasoning LLMs. With in-context experiences of failure, some but not all LLMs reduce their overconfidence leading to significantly improved decision making, while others do not. Interestingly, all LLMs' decisions are approximately rational given their estimated probabilities of success, yet their overly-optimistic estimates result in poor decision making. These results suggest that current LLM agents are hindered by their lack of awareness of their own capabilities. We discuss the implications of LLMs' awareness of their capabilities for AI misuse and misalignment risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。