修复错误假设会损害模型对正常问题的回答能力
Don't `Well, Actually' Me Unless You Know What You're Talking About: Weak Presupposition Verification Degrades General QA Performance

- 通过提取并验证前提来检测错误假设
- 越擅长处理错误假设,越可能误判正确前提
- 适合关注真实场景下模型泛化能力的研究者
虚假前提问答(FPQA)测试大模型识别问题中错误假设的能力,并要求其避免强化错误认知或主动修正。当前方法通常将任务拆解为提取前提并逐一验证。尽管专用基准测试性能持续提升,但评估主要聚焦于含错误假设的问题(FPQ),忽视了对正常问题(TPQ)的表现。由于多数基准中FPQ占比远高于自然语料,导致评测结果无法反映真实场景下的表现。我们对多种模型族、规模及基准进行广泛实验,发现更擅长处理FPQ的模型在TPQ上表现反而更差。分析表明,这是由于事实核查模块过于敏感,会错误拒绝真实前提。我们希望这些发现能引导未来研究设计更具泛化能力的FPQA方法。
原文摘要 · Abstract (English)
False-presupposition QA (FPQA) tests LLMs on their ability to identify false presuppositions in questions and abstain or correct them rather than reinforcing false assumptions. The common approach reduces the task to prompting LLMs to extract presuppositions and fact checking each presupposition. While the performance on dedicated benchmarks keeps improving, evaluation largely focuses on questions with false presuppositions (FPQs) while ignoring the performance on ``normal'' questions (TPQs). Since many benchmarks over-represent FPQs compared to their natural occurrence, the result is that performance on these benchmarks doesn't reflect real-world QA performance. Through extensive experiments across various model families, sizes, and benchmarks, we show that methods that perform better on FPQs tend to perform worse on TPQs. Our analysis reveals this is the result of weak fact checking modules that reject also true presuppositions. We hope our findings will help guide future work toward FPQA methods that generalize well to realistic settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。