通过显式诊断错误提升问答生成评估准确性
ErrEval: Error-Aware Evaluation for Question Generation through Explicit Diagnostics
- 分两阶段评估:先诊断错误类型,再基于诊断结果打分
- 在三个数据集上显著提升与人工判断的一致性
- 适合关注生成质量、需避免高估低质问题的研究者
自动问答生成常出现事实幻觉和答案不匹配等严重缺陷。现有评估方法(包括基于大模型的评估)多采用黑箱整体评分,缺乏显式的错误建模,导致忽略缺陷并高估问题质量。为此,我们提出ErrEval——一种灵活且错误感知的评估框架,通过显式错误诊断增强问答生成评估。具体地,将评估重构为两阶段:先由轻量级错误识别器检测并分类结构、语言和内容层面的常见错误;再将诊断信号作为显式证据,引导大模型评估者做出更细粒度、更可信的判断。在三个基准上的大量实验表明,引入显式诊断能显著提升与人工判断的一致性,有效缓解对低质量问题的高估。
原文摘要 · Abstract (English)
Automatic Question Generation (QG) often produces outputs with critical defects, such as factual hallucinations and answer mismatches. However, existing evaluation methods, including LLM-based evaluators, mainly adopt a black-box and holistic paradigm without explicit error modeling, leading to the neglect of such defects and overestimation of question quality. To address this issue, we propose ErrEval, a flexible and Error-aware Evaluation framework that enhances QG evaluation through explicit error diagnostics. Specifically, ErrEval reformulates evaluation as a two-stage process of error diagnosis followed by informed scoring. At the first stage, a lightweight plug-and-play Error Identifier detects and categorizes common errors across structural, linguistic, and content-related aspects. These diagnostic signals are then incorporated as explicit evidence to guide LLM evaluators toward more fine-grained and grounded judgments. Extensive experiments on three benchmarks demonstrate the effectiveness of ErrEval, showing that incorporating explicit diagnostics improves alignment with human judgments. Further analyses confirm that ErrEval effectively mitigates the overestimation of low-quality questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。