arXiv:2503.02863cs.CLcs.LG2025-03NeurIPS被引 21

通过提示词引导让大模型更准确地表达信心,提升可靠性。

SteerConf: Steering LLMs for Confidence Elicitation

  • 用不同强度的提示词引导模型输出特定方向的信心值
  • 多方向信心一致性度量使校准效果提升显著
  • 无需微调即可适配现有大模型,适合安全应用

大型语言模型在多个领域表现优异,但常出现过度自信问题,限制其在关键场景中的可靠性。本文提出SteerConf框架,系统性地引导模型输出更准确的信心分数。该框架包含三个核心组件:(1) 提示词引导策略,通过具有不同引导强度的提示词,促使模型向保守或乐观方向输出信心值;(2) 修正后信心一致性度量,量化多个引导方向下的信心一致性,增强校准效果;(3) 修正信心校准方法,结合一致性度量对信心分数进行线性量化聚合以选择答案。该方法无需额外训练或微调,适用于现有大模型。在涵盖专业知识、常识、伦理和推理任务的七个基准上,使用GPT-3.5、LLaMA 3、GPT-4等先进模型进行实验,结果表明SteerConf显著优于现有方法,多数情况下提升明显。研究揭示了通过引导模型信心来提升其可靠性、实现更安全部署的巨大潜力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) exhibit impressive performance across diverse domains but often suffer from overconfidence, limiting their reliability in critical applications. We propose SteerConf, a novel framework that systematically steers LLMs' confidence scores to improve their calibration and reliability. SteerConf introduces three key components: (1) a steering prompt strategy that guides LLMs to produce confidence scores in specified directions (e.g., conservative or optimistic) by leveraging prompts with varying steering levels; (2) a steered confidence consistency measure that quantifies alignment across multiple steered confidences to enhance calibration; and (3) a steered confidence calibration method that aggregates confidence scores using consistency measures and applies linear quantization for answer selection. SteerConf operates without additional training or fine-tuning, making it broadly applicable to existing LLMs. Experiments on seven benchmarks spanning professional knowledge, common sense, ethics, and reasoning tasks, using advanced LLM models (GPT-3.5, LLaMA 3, GPT-4), demonstrate that SteerConf significantly outperforms existing methods, often by a significant margin. Our findings highlight the potential of steering the confidence of LLMs to enhance their reliability for safer deployment in real-world applications.

大模型信心校准提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。