用大模型比较法无真值估算题目难度,高效且可扩展。
Estimating problem difficulty without ground truth using Large Language Model comparisons
- 让大模型两两比较题目难易,用布拉德利-特里模型计算得分。
- 与人类标注相关性高达0.80(n=1876),且抗幻觉干扰能力好。
- 无需真值、可跨分布适用,适合模型评估和智能研究辅助。
大语言模型微调的进展显著提升了其在基准测试中的表现,推动了对更复杂合成数据的需求。生成此类数据的关键步骤是估算题目难度。现有方法如人工校准或基于性能的评分,因依赖真值、难以扩展且无法泛化到人类和大模型均无法解决的分布外问题,面临局限。为此,本文提出新方法LLM compare:让大模型进行两两难度比较,再基于结果计算布拉德利-特里分数。我们构建了一个三维框架,将现有方法定位在构造、规模与依赖性三个维度上,发现LLM compare自然占据所有理想象限,是首个连续动态、模型无关且不依赖真值的度量方法。验证显示,其与人类标注的相关性达皮尔逊$ r \≥q 0.80 $(n=1876);在注入10%噪声情况下,相关性下降不足6%。本工作为替代耗时的人工标注与合成数据生成提供了可能,将在课程设计、模型评估与人工智能辅助研究中发挥重要作用。
原文摘要 · Abstract (English)
Recent advances in the finetuning of large language models (LLMs) have significantly improved their performance on established benchmarks, emphasizing the need for increasingly difficult, synthetic data. A key step in this data generation pipeline is a method for estimating problem difficulty. Current approaches, such as human calibration or performance-based scoring, fail to generalize to out-of-distribution problems, i.e. problems currently unsolvable by humans and LLMs, because they are not scalable, time-consuming, and ground truth dependent. Therefore, we propose a new method for estimating problem difficulty, LLM compare, that addresses these limitations. An LLM performs pairwise difficulty comparisons, and then Bradley-Terry scores are computed based on the outcomes. To validate our method, we first propose a conceptual framework that positions existing approaches on three orthogonal planes--construction, scale and dependence--identifying which quadrants a measure needs to occupy to score out-of-distribution problems. LLM compare naturally occupies all desirable quadrants as the first measure that is continuous and dynamic, model-agnostic and independent of ground truth information. As a second validation, we show that LLM compare demonstrates strong alignment with human annotations: Pearson $r \geq 0.80$ for $n=1876$. Thirdly, we show that LLM compare is robust to hallucinations, with less than $6\%$ degradation in Pearson correlation for $10\%$ noise injection. Our work represents a significant step towards replacing time-consuming human annotations and synthetic data generation, and will be an important driver for curriculum design, model evaluation, and AI-assisted research ideation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。