arXiv:2508.10022cs.CLcs.AI2025-08被引 1

用统计检验提升大模型答题可信度,确保错误率可控。

Conformal P-Value in Multiple-Choice Question Answering Tasks with Provable Risk Control

  • 结合p值与自一致性采样,生成可验证的答题置信区间。
  • 在MMLU和MMLU-Pro上实现用户指定的错误率,且置信集大小随风险上升而减小。
  • 适合高风险问答场景,如医疗、法律等需严格可靠性保障的应用。

本研究提出一种融合显著性检验的分位数预测(CP)框架,以增强大语言模型在多选题问答任务中的可信度。尽管大模型已在学科问答中广泛应用,但幻觉与非事实生成严重损害了回答可靠性。虽然CP能提供严格的边缘覆盖保证,显著性检验也具备成熟的统计严谨性,但两者协同机制尚未探索。为缓解幻觉与事实错误,该框架通过自一致性重采样计算选项频率,克服大模型黑箱特性,基于零假设检验(H₀)和经验p值构建预测集。在MMLU和MMLU-Pro基准上使用现成大模型评估显示:(1) 增强型CP实现了用户设定的实证误覆盖率;(2) 测试集平均预测集大小(APSS)随风险水平α单调下降,验证其作为不确定性度量的有效性。本工作建立了高风险问答应用中可信大模型部署的原理性统计框架。

原文摘要 · Abstract (English)

This study introduces a significance testing-enhanced conformal prediction (CP) framework to improve trustworthiness of large language models (LLMs) in multiple-choice question answering (MCQA). While LLMs have been increasingly deployed in disciplinary QA scenarios, hallucination and nonfactual generation substantially compromise response reliability. Although CP provides statistically rigorous marginal coverage guarantees for prediction sets, and significance testing offers established statistical rigor, their synergistic integration remains unexplored. To mitigate hallucination and factual inaccuracies, our framework integrates $p$-value computation with conformity scoring through self-consistency resampling of MCQA responses. This approach calculates option frequencies to address LLMs' black-box nature, subsequently constructing prediction sets via null hypothesis testing ($\mathcal{H}_0$) with empirically derived $p$-values. Evaluations on MMLU and MMLU-Pro benchmarks using off-the-shelf LLMs demonstrate: (1) The enhanced CP achieves user-specified empirical miscoverage rates; (2) Test-set average prediction set size (APSS) decreases monotonically with increasing risk levels ($α$), validating APSS as an effective uncertainty metric. This work establishes a principled statistical framework for trustworthy LLM deployment in high-stakes QA applications.

大模型可信度统计推断多选题问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。