提出新方法提升开放问答中答案集的可靠性,减少幻觉。
MiRD: Reliable Set-Valued Prediction for Open-Ended Question Answering via Miscoverage Risk Decomposition

- 分解误覆盖风险为采样与选择两阶段,更精准控制不确定性。
- 在三个数据集上验证,显著降低整体误覆盖率,预测集更紧凑。
- 适合追求高可靠性开放问答系统的研究者与开发者。
可靠集合值预测为缓解开放问答中的幻觉问题提供了理论基础,但现有基于校准的方法通常依赖一个脆弱假设:有限采样必须已产生至少一个可接受的答案,否则校准样本将被丢弃。本文提出MiRD,一种两阶段框架,将总体误覆盖分解为采样失败和条件选择失败。第一阶段,MiRD建立在固定预算下,有限采样未产生任何可接受答案的概率的期望水平边际上界。第二阶段,在采样成功条件下,利用全校准集上的与准入相关的非同质性分数校准一致选择阈值,从而保持校准集完整性。在三个开放问答数据集和八种模型上,MiRD有效控制了采样风险、条件选择风险和总体误覆盖,且第一阶段边界比PAC类方法更紧,预测集比仅成功校准更具适应性。
原文摘要 · Abstract (English)
Reliable set-valued prediction provides a principled way to mitigate hallucinations in open-ended question answering (QA), yet existing conformal approaches typically rely on a fragile premise: finite sampling must already produce at least one admissible candidate, or calibration examples violating this condition are discarded. In this paper, we introduce MiRD, a two-stage framework that decomposes overall miscoverage into sampling failure and conditional selection failure. In Stage I, MiRD establishes an expectation-level marginal upper bound on the probability that finite sampling produces no admissible answer under a fixed budget. In Stage II, conditioned on sampling success, MiRD calibrates a conformal selection threshold using admission-correlated nonconformity scores defined over the full calibration set, thereby preserving calibration-set integrity. Across three open-ended QA datasets and eight models, MiRD controls sampling risk, conditional selection risk, and overall miscoverage, while yielding tighter first-stage bounds than PAC-style alternatives and more adaptive prediction sets than successful-only calibration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。