提升医学多选题问答的可信度,用统计方法控制错误率。
Correctness Coverage Evaluation for Medical Multiple-Choice Question Answering Based on the Enhanced Conformal Prediction Framework
- 结合正确选项频次与自一致性,增强置信区间评估
- 在三个数据集上实现指定错误率保证,平均预测集更小
- 适合医疗AI风险敏感场景,提升模型可靠性
大语言模型(LLMs)在医学问答中应用日益广泛,但其易产生幻觉和非事实信息,影响高风险医疗任务中的可信度。置信预测(CP)提供严格的边缘覆盖率保障,但在医学问答中研究有限。本文提出一种增强型CP框架,用于医学多选题问答(MCQA)。通过将非符合度分数与正确选项频率关联,并利用自一致性机制,解决模型内部不透明问题,引入单调损失函数的风险控制策略。在MedMCQA、MedQA和MMLU数据集上,使用四个现成的LLM进行评估,所提方法在提高风险水平时仍满足指定错误率约束,同时降低平均预测集大小,为大模型不确定性评估提供了有前景的指标。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly adopted in medical question-answering (QA) scenarios. However, LLMs can generate hallucinations and nonfactual information, undermining their trustworthiness in high-stakes medical tasks. Conformal Prediction (CP) provides a statistically rigorous framework for marginal (average) coverage guarantees but has limited exploration in medical QA. This paper proposes an enhanced CP framework for medical multiple-choice question-answering (MCQA) tasks. By associating the non-conformance score with the frequency score of correct options and leveraging self-consistency, the framework addresses internal model opacity and incorporates a risk control strategy with a monotonic loss function. Evaluated on MedMCQA, MedQA, and MMLU datasets using four off-the-shelf LLMs, the proposed method meets specified error rate guarantees while reducing average prediction set size with increased risk level, offering a promising uncertainty evaluation metric for LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。