arXiv:2409.18892cs.CL2024-09NeurIPS被引 6

用考试题区分能力的思路,让大模型评测更精准。

IDGen: Item Discrimination Induced Prompt Generation for LLM Evaluation

  • 基于题目区分度理论生成能分辨模型优劣的评测题
  • 新数据集平均分51.92,方差10.06,区分度远超旧方法
  • 适合想做严谨模型对比或评测基准研究的人

随着大语言模型(LLMs)在复杂任务上的表现日益提升,评测集必须同步进化以保持足够的区分能力。借鉴教育评估中的项目区分度(ID)理论,本文提出一种受ID启发的提示生成框架,使评测集能动态适应模型能力。该数据合成框架兼顾广度与针对性,可生成全面评估模型能力并揭示模型间性能差异的提示,有效区分不同模型在多任务、多领域中的优劣。为保证数据质量,引入自修正机制,并构建两个模型分别预测提示的区分度与难度,助力高质量数据生成。将生成数据用于评估五种SOTA模型,结果平均得分为51.92,方差达10.06;而此前方法(SELF-INSTRUCT和WizardLM)平均分超67,方差低于3.2。实验表明,本框架生成的数据更具挑战性与区分度。我们将发布超过3,000个精心设计的提示构成的数据集,以推动大模型评测研究。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) grow increasingly adept at managing complex tasks, the evaluation set must keep pace with these advancements to ensure it remains sufficiently discriminative. Item Discrimination (ID) theory, which is widely used in educational assessment, measures the ability of individual test items to differentiate between high and low performers. Inspired by this theory, we propose an ID-induced prompt synthesis framework for evaluating LLMs to ensure the evaluation set can continually update and refine according to model abilities. Our data synthesis framework prioritizes both breadth and specificity. It can generate prompts that comprehensively evaluate the capabilities of LLMs while revealing meaningful performance differences between models, allowing for effective discrimination of their relative strengths and weaknesses across various tasks and domains. To produce high-quality data, we incorporate a self-correct mechanism into our generalization framework, and develop two models to predict prompt discrimination and difficulty score to facilitate our data synthesis framework, contributing valuable tools to evaluation data synthesis research. We apply our generated data to evaluate five SOTA models. Our data achieves an average score of 51.92, accompanied by a variance of 10.06. By contrast, previous works (i.e., SELF-INSTRUCT and WizardLM) obtain an average score exceeding 67, with a variance below 3.2. The results demonstrate that the data generated by our framework is more challenging and discriminative compared to previous works. We will release a dataset of over 3,000 carefully crafted prompts to facilitate evaluation research of LLMs.

大模型评测提示生成区分度数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。