用答案可信度熵值评估问题难度,更准且更稳。
Question Difficulty Estimation for Large Language Models via Answer Plausibility Scoring

- 通过候选答案的可信度得分熵值衡量问题难度。
- 在4个数据集上均优于现有方法,且对参数变化不敏感。
- 结果与人评难度高度一致,适合评测和优化大模型。
评估问题难度是提升大型语言模型问答能力的关键。现有方法多依赖可读性公式、检索信号或流行度统计,难以捕捉现代大模型面临的推理挑战。本文提出Q-DAPS(基于答案可信度评分的问题难度估计),通过计算候选答案可信度得分的熵值来评估难度。在TriviaQA、NQ、MuSiQue和QASC四个主流QA数据集上系统验证,Q-DAPS表现持续优于基线。该方法对超参数变化和问题类型具有强鲁棒性,消融实验显示其在不同可信度估计范式、模型规模及真实场景下均稳定有效。人工评估进一步证实,Q-DAPS的难度估计与人类判断高度一致。整体上,Q-DAPS提供了一种可解释、可扩展且抗偏差的问题难度评估方法。
原文摘要 · Abstract (English)
Estimating question difficulty is a critical component in evaluating and improving large language models (LLMs) for question answering (QA). Existing approaches often rely on readability formulas, retrieval-based signals, or popularity statistics, which may not fully capture the reasoning challenges posed to modern LLMs. In this paper, we introduce Q-DAPS (Question Difficulty based on Answer Plausibility Scores) method, a novel approach that estimates question difficulty by computing the entropy of plausibility scores over candidate answers. We systematically evaluate Q-DAPS across four prominent QA datasets-TriviaQA, NQ, MuSiQue, and QASC-demonstrating that it consistently outperforms baselines. Moreover, Q-DAPS shows strong robustness across hyperparameter variations and question types. Extensive ablation studies further show that Q-DAPS remains robust across different plausibility estimation paradigms, model sizes, and realistic settings. Human evaluations further confirm strong alignment between Q-DAPS's difficulty estimates and human judgments of question difficulty. Overall, Q-DAPS provides an interpretable, scalable, and bias-resilient approach to question difficulty estimation in modern QA systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。