梳理大模型诚实性问题,明确评估与改进方法。
A Survey on the Honesty of Large Language Models
- 从认知边界出发定义模型诚实性,区分已知与未知知识。
- 总结现有评估方法,揭示模型自信说错话的普遍现象。
- 适合关注AI对齐、可信生成的研究者与从业者参考。
诚实是大语言模型(LLMs)与人类价值观对齐的基本原则,要求模型能够识别自身已知与未知信息,并如实表达知识。尽管前景可观,当前大模型仍存在显著不诚实行为,如自信地给出错误答案或无法表达已知内容。此外,相关研究面临定义模糊、难区分知识状态、整体理解不足等挑战。为此,本文提供关于大模型诚实性的综述,涵盖其概念澄清、评估方法及改进策略,并为未来研究提供洞见,旨在推动该重要领域的发展。
原文摘要 · Abstract (English)
Honesty is a fundamental principle for aligning large language models (LLMs) with human values, requiring these models to recognize what they know and don't know and be able to faithfully express their knowledge. Despite promising, current LLMs still exhibit significant dishonest behaviors, such as confidently presenting wrong answers or failing to express what they know. In addition, research on the honesty of LLMs also faces challenges, including varying definitions of honesty, difficulties in distinguishing between known and unknown knowledge, and a lack of comprehensive understanding of related research. To address these issues, we provide a survey on the honesty of LLMs, covering its clarification, evaluation approaches, and strategies for improvement. Moreover, we offer insights for future research, aiming to inspire further exploration in this important area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。