用熵值加正确性探测,让大模型更可靠地拒绝高风险回答。
Entropy Alone is Insufficient for Safe Selective Prediction in LLMs
- 结合熵值与正确性探测信号,提升拒绝决策准确性。
- 在三个问答数据集上,错误率降低15%以上,覆盖率达80%以上。
- 适合关注模型安全部署、需低错误率场景的研究者使用。
选择性预测系统可通过在高风险情况下拒答来缓解大语言模型幻觉带来的危害。尽管不确定性量化技术常用于识别此类情况,但其在更广泛的预测策略中,特别是在实现低目标错误率方面的表现却很少被评估。本文揭示了基于熵的不确定性方法存在一种依赖模型的失效模式,导致拒答行为不可靠,并提出通过将熵值分数与正确性探测信号结合来解决该问题。在三个问答基准(TriviaQA、BioASQ、MedicalQA)和四个模型家族上,联合评分显著优于仅使用熵的基线,在风险-覆盖率权衡和校准性能方面均有提升。结果强调了面向部署的不确定性评估的重要性,应采用直接反映系统能否在指定风险水平下可信运行的指标。
原文摘要 · Abstract (English)
Selective prediction systems can mitigate harms resulting from language model hallucinations by abstaining from answering in high-risk cases. Uncertainty quantification techniques are often employed to identify such cases, but are rarely evaluated in the context of the wider selective prediction policy and its ability to operate at low target error rates. We identify a model-dependent failure mode of entropy-based uncertainty methods that leads to unreliable abstention behaviour, and address it by combining entropy scores with a correctness probe signal. We find that across three QA benchmarks (TriviaQA, BioASQ, MedicalQA) and four model families, the combined score generally improves both the risk--coverage trade-off and calibration performance relative to entropy-only baselines. Our results highlight the importance of deployment-facing evaluation of uncertainty methods, using metrics that directly reflect whether a system can be trusted to operate at a stated risk level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。