发现问答题模糊导致模型出错,重写后性能显著提升
Who is the richest club in the championship? Detecting and Rewriting Underspecified Questions Improve QA Performance
- 用大模型识别问题模糊性,发现16%到50%的问题不明确
- 重写模糊问题后,问答准确率明显提高
- 适合关注评测公平性和问题设计的研究者
大型语言模型在表述清晰的问题上表现良好,但标准问答基准仍远未解决。我们认为,这一差距部分源于问题的不明确性——即缺乏额外上下文时无法唯一确定其含义。为此,我们引入基于大模型的分类器来识别不明确问题,并应用于多个常用问答数据集,发现16%至超过50%的基准问题存在不明确性,且大模型在这些问题上的表现显著下降。为隔离不明确性的影响,我们进行了一项受控重写实验:将不明确问题改写为完整明确的形式,同时保持正确答案不变。在此设定下,问答性能持续提升,表明许多看似模型能力不足的失败实则源于问题本身不明确。研究结果强调了不明确性作为问答评估中重要混淆因素的作用,也呼吁在基准设计中更加重视问题清晰度。
原文摘要 · Abstract (English)
Large language models (LLMs) perform well on well-posed questions, yet standard question-answering (QA) benchmarks remain far from solved. We argue that this gap is partly due to underspecified questions - queries whose interpretation cannot be uniquely determined without additional context. To test this hypothesis, we introduce an LLM-based classifier to identify underspecified questions and apply it to several widely used QA datasets, finding that 16% to over 50% of benchmark questions are underspecified and that LLMs perform significantly worse on them. To isolate the effect of underspecification, we conduct a controlled rewriting experiment that serves as an upper-bound analysis, rewriting underspecified questions into fully specified variants while holding gold answers fixed. QA performance consistently improves under this setting, indicating that many apparent QA failures stem from question underspecification rather than model limitations. Our findings highlight underspecification as an important confound in QA evaluation and motivate greater attention to question clarity in benchmark design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。