arXiv:2409.00352cs.CLcs.LG2024-09被引 1

研究发现对齐训练会损害大模型置信度校准能力

Does Alignment Tuning Really Break LLMs' Internal Confidence?

  • 从四个维度系统分析对齐如何影响模型置信度
  • 严格条件下对齐均导致校准性能下降
  • 适合关注大模型可靠性与可信度的研究者

大语言模型在实际应用中需要可靠的置信度校准。本研究从模型、校准指标、任务和置信度提取方法四个维度,全面分析了大模型校准退化的现象。初步分析显示对齐与校准并非总是权衡关系,但在更严格的分析条件下,对齐过程始终损害校准性能。这表明需谨慎设计置信度测量方式,并推动能同时实现指令遵循与良好校准的算法研究。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown remarkable progress, but their real-world application necessitates reliable calibration. This study conducts a comprehensive analysis of calibration degradation of LLMs across four dimensions: models, calibration metrics, tasks, and confidence extraction methods. Initial analysis showed that the relationship between alignment and calibration is not always a trade-off, but under stricter analysis conditions, we found the alignment process consistently harms calibration. This highlights the need for (1) a careful approach when measuring model confidences and calibration errors and (2) future research into algorithms that can help LLMs to achieve both instruction-following and calibration without sacrificing either.

大模型校准置信度对齐训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。