用动态生成题库对抗大模型评估中的数据污染问题
AdEval: Alignment-based Dynamic Evaluation to Mitigate Data Contamination in Large Language Models
- 基于知识对齐动态构建评测题,避免直接使用污染数据
- 六维认知层级出题,提升评估复杂度与多样性
- 适合关注评测公平性的大模型研究人员
大型语言模型(LLMs)在超大规模语料上预训练,数据污染问题日益严重,静态评测基准可能高估模型性能。为此,本文提出一种名为AdEval(基于对齐的动态评估)的动态评估方法。AdEval首先从静态数据集提取知识点与核心思想,实现与静态基准的核心内容动态对齐,通过不直接依赖静态数据集,从源头降低数据污染风险。随后,通过在线搜索获取背景信息,生成知识点的详细描述。最后,依据布卢姆认知分类法,在记忆、理解、应用、分析、评价、创造六个维度设计问题,实现多层级认知评估。此外,通过迭代重构问题控制动态生成数据集的复杂度。在多个数据集上的实验结果表明,AdEval有效缓解数据污染对评估结果的影响,解决了复杂度控制不足与评估维度单一的问题,提升了大模型评估的公平性、可靠性与多样性。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) are pre-trained on ultra-large-scale corpora, the problem of data contamination is becoming increasingly serious, and there is a risk that static evaluation benchmarks overestimate the performance of LLMs. To address this, this paper proposes a dynamic data evaluation method called AdEval (Alignment-based Dynamic Evaluation). AdEval first extracts knowledge points and main ideas from static datasets to achieve dynamic alignment with the core content of static benchmarks, and by avoiding direct reliance on static datasets, it inherently reduces the risk of data contamination from the source. It then obtains background information through online searches to generate detailed descriptions of the knowledge points. Finally, it designs questions based on Bloom's cognitive hierarchy across six dimensions-remembering, understanding, applying, analyzing, evaluating, and creating to enable multi-level cognitive assessment. Additionally, AdEval controls the complexity of dynamically generated datasets through iterative question reconstruction. Experimental results on multiple datasets show that AdEval effectively alleviates the impact of data contamination on evaluation results, solves the problems of insufficient complexity control and single-dimensional evaluation, and improves the fairness, reliability and diversity of LLMs evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。