自动生成高质量干扰项,让生成模型评估更可靠。
D-GEN: Automatic Distractor Generation and Evaluation for Reliable Assessment of Generative Model
- 用生成模型将开放回答转为选择题格式。
- 生成干扰项与真实干扰项排名一致(Spearman相关0.99)。
- 适合需要高效评估生成模型的研究者使用。
开放式生成的模型评估因回答格式不一致而困难。多选题评估可缓解此问题,但高质量干扰项的生成耗时费力。我们提出D-GEN,首个开源的干扰项生成模型,能将开放式数据转化为多选题格式。为评估干扰项质量,我们提出两种新方法:(1) 排名一致性,确保生成干扰项保留真实干扰项的区分能力;(2) 熵分析,比较模型置信度分布。结果表明,D-GEN保持了高度的排名一致性(斯皮尔曼相关系数0.99,肯德尔等级相关0.94),且熵分布接近真实干扰项。人工评估确认生成内容流畅、连贯、具有迷惑性且错误合理。本工作推动了自动化、可靠的干扰项生成与评估,为多选题评估树立了新标准。
原文摘要 · Abstract (English)
Evaluating generative models with open-ended generation is challenging due to inconsistencies in response formats. Multiple-choice (MC) evaluation mitigates this issue, but generating high-quality distractors is time-consuming and labor-intensive. We introduce D-GEN, the first open-source distractor generator model that transforms open-ended data into an MC format. To evaluate distractor quality, we propose two novel methods: (1) ranking alignment, ensuring generated distractors retain the discriminatory power of ground-truth distractors, and (2) entropy analysis, comparing model confidence distributions. Our results show that D-GEN preserves ranking consistency (Spearman's rho 0.99, Kendall's tau 0.94) and closely matches the entropy distribution of ground-truth distractors. Human evaluation further confirms the fluency, coherence, distractiveness, and incorrectness. Our work advances robust and efficient distractor generation with automated evaluation, setting a new standard for MC evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。