arXiv:2503.11256cs.CLcs.AI2025-03中稿 · NAACL被引 8

通过让大模型自主设定回答边界,发现其自我认知可信度不足80%。

Line of Duty: Evaluating LLM Self-Knowledge via Consistency in Feasibility Boundaries

  • 让模型自己定义可回答范围,分析其一致性来评估自我认知能力。
  • GPT-4o等前沿模型在超80%情况下不确定自身能力,信任度低。
  • 模型在时间感知和上下文理解上最弱,适合研究自我认知的学者参考。

随着大语言模型能力增强,其最深刻成就可能是知道何时说‘我不知道’。现有研究受限于人类定义的可行性标准,忽视了模型无法回答的原因,也未深入分析自我认知的缺陷类型。本研究提出新方法:允许模型自主设定可行性边界,并分析其边界的一致性。结果发现,即使前沿模型如GPT-4o和Mistral Large,在超过80%的情况下也不确定自身能力,显示出显著的信任问题。分析表明,模型在不同任务类别间在过度自信与保守之间摇摆,最严重的自我认知弱点在于时间感知和上下文理解。这些上下文理解困难导致模型质疑自身操作边界,引发严重自我认知混乱。代码与结果已公开于https://github.com/knowledge-verse-ai/LLM-Self_Knowledge_Eval。

原文摘要 · Abstract (English)

As LLMs grow more powerful, their most profound achievement may be recognising when to say "I don't know". Existing studies on LLM self-knowledge have been largely constrained by human-defined notions of feasibility, often neglecting the reasons behind unanswerability by LLMs and failing to study deficient types of self-knowledge. This study aims to obtain intrinsic insights into different types of LLM self-knowledge with a novel methodology: allowing them the flexibility to set their own feasibility boundaries and then analysing the consistency of these limits. We find that even frontier models like GPT-4o and Mistral Large are not sure of their own capabilities more than 80% of the time, highlighting a significant lack of trustworthiness in responses. Our analysis of confidence balance in LLMs indicates that models swing between overconfidence and conservatism in feasibility boundaries depending on task categories and that the most significant self-knowledge weaknesses lie in temporal awareness and contextual understanding. These difficulties in contextual comprehension additionally lead models to question their operational boundaries, resulting in considerable confusion within the self-knowledge of LLMs. We make our code and results available publicly at https://github.com/knowledge-verse-ai/LLM-Self_Knowledge_Eval

大模型自我认知可信度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。