用心理测量学方法更精准评估大模型的定理证明能力,减少测试量并凸显性能差异。
Psychometric-Based Evaluation for Theorem Proving with Large Language Models
- 基于心理测量学为定理分配难度与区分度标签,构建graded数据集
- 实测显示新方法仅用23%题目即能分辨模型真实差距
- 适合需要高效、公平对比大模型证明能力的研究者使用
用于形式化定理证明的大语言模型(LLMs)已成为研究热点。当前主要通过miniF2F等数据集上的证明通过率评估其能力,但该方法忽视了定理重要性差异,导致无法真实反映模型间性能差距,并带来高评估成本。本文提出一种基于心理测量学的评估方法,包含数据集标注与自适应评估两部分。首先,设计指标计算方法,对miniF2F数据集中每个定理进行难度与区分度标注,生成增强版数据集miniF2F-Graded。实验表明,该标注能更好反映LLM感知的定理难度。其次,设计自适应评估机制,根据标注结果和模型实时表现动态选择最适定理进行测试。应用于10个LLM的评估中,结果表明该方法可精细揭示模型间性能差异,且仅需原数据集23%的题目即可完成有效评估。
原文摘要 · Abstract (English)
Large language models (LLMs) for formal theorem proving have become a prominent research focus. At present, the proving ability of these LLMs is mainly evaluated through proof pass rates on datasets such as miniF2F. However, this evaluation method overlooks the varying importance of theorems. As a result, it fails to highlight the real performance disparities between LLMs and leads to high evaluation costs. This study proposes a psychometric-based evaluation method for theorem proving with LLMs, comprising two main components: Dataset Annotation and Adaptive Evaluation. First, we propose a metric calculation method to annotate the dataset with difficulty and discrimination metrics. Specifically, we annotate each theorem in the miniF2F dataset and grade them into varying difficulty levels according to the performance of LLMs, resulting in an enhanced dataset: miniF2F-Graded. Experimental results show that the difficulty grading in miniF2F-Graded better reflects the theorem difficulty perceived by LLMs. Secondly, we design an adaptive evaluation method to dynamically select the most suitable theorems for testing based on the annotated metrics and the real-time performance of LLMs. We apply this method to evaluate 10 LLMs. The results show that our method finely highlights the performance disparities between LLMs. It also reduces evaluation costs by using only 23% of the theorems in the dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。