让大模型学会评估自己知道什么,提升自我认知能力。
Fine-Tuning Language Models to Know What They Know
- 用 $d'_{\rm type2}$ 量化模型的元认知能力,排除偏见干扰。
- 新方法在未见数据、语言和新知识上均表现稳定,泛化性强。
- 只需调整少量参数即可提升元认知,便于精准优化。
由于存在偏见和启发式策略,评估大语言模型(LLMs)的真实元认知能力十分困难。本文提出一个框架,用于测量并增强LLM的元认知能力,同时控制这些偏见。通过 $d'_{\rm type2}$ 指标建立测量方法,可有效分离出元认知能力。提出的元认知对齐进化策略(ESMA)在未见过的数据集、语言及新增知识上展现出稳健的泛化能力。最后的参数分析表明,这些改进由一组稀疏参数驱动,为针对性的元认知优化提供了新路径。
原文摘要 · Abstract (English)
Evaluating true metacognition in Large Language Models (LLMs) is difficult due to biases and heuristics. This paper presents a framework to measure and enhance LLM metacognition while controlling for these biases. A measurement method using the $d'_{\rm type2}$ metric is established to isolate metacognitive ability. The Evolution Strategy for Metacognitive Alignment (ESMA) is proposed, demonstrating robust generalization across unseen datasets, languages, and newly acquired knowledge. Finally, parameter analysis reveals that these improvements are driven by a sparse set of parameters, offering new pathways for targeted metacognitive optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。