优化大模型答案审查顺序,提升有限人力下纠错效率。
Risk Is Not Review Value: Wrong-Answer Exposure Under Bounded Review Budgets

- 用审查价值综合错误率、可修复性、影响和成本来排序待审答案。
- 在720项压力测试中,审查价值排序使修复后仍暴露的错误率从88.1%降至71.6%。
- 适合关注大模型可信评估与资源受限审查场景的研究者。
大型语言模型助手生成的答案常超出人工审查能力。传统评估关注答案是否错误或无依据,但受限审查预算时需决定优先审查哪些答案。仅看风险不足:高风险答案可能难以修复,中等风险答案却可能直接通过已有证据修正。本文将审查优先级建模为暴露度降低,审查价值融合错误估计、干预可行性、影响与成本。采用错误答案暴露率(WAER)与修复后残余暴露率(PRRE)评估审查队列。在720项的TAT-QA/SciFact压力基准上,审查价值排序在20%审查预算下保持答案级WAER为0.600(原0.605),但使修复后残余暴露率从0.881降至0.716。结果表明,可信评估需兼顾错误检测与有限审查下暴露错误的减少。
原文摘要 · Abstract (English)
LLM assistants often produce more answers than humans can review before users see them. Most evaluations ask whether an answer is wrong, unsupported, or low-confidence. Bounded review budgets instead ask which answers should be checked first under a fixed review budget. Risk alone is not enough: a high-risk answer may be hard to repair, while a moderately risky answer may be directly correctable from available evidence. For generated-answer evaluation, we model review prioritization as exposure reduction, where review value combines estimated wrongness, intervention affordance, impact, and cost. We evaluate review queues with Wrong-Answer Exposure Ratio (WAER), the fraction of wrong answers left unreviewed, and post-repair residual exposure (PRRE), the fraction still exposed after deterministic benchmark-supported repairs. PRRE uses repairability rules that do not numerically reuse the affordance scores used for ranking. On a 720-item TAT-QA/SciFact stress benchmark, review-value ranking keeps answer-level WAER nearly unchanged at 20% budget (0.605 vs. 0.600) but lowers PRRE from 0.881 to 0.716. These results show that trustworthy LLM evaluation should measure not only error detection, but also how limited review capacity reduces exposed wrong answers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。