用户只需提供文档,即可自动生成精准、低成本的LLM评估集。
YourBench: Easy Custom Evaluation Sets for Everyone
- 输入文档自动构建动态评估集,无需人工标注。
- 生成评估集成本低于15美元,且模型排名与原基准完全一致。
- 适合需要快速定制领域评估的开发者和研究者。
高效评估大语言模型仍是关键瓶颈,传统静态基准易饱和污染,人工评估成本高且耗时。本文提出 YourBench,一个开源框架,可从用户提供的文档中自动、动态生成可靠、及时、领域定制的评估集,成本低且无需人工标注。通过仅用少量源文本复现了7个MMLU子集,总推理成本不足15美元,模型性能排序保持不变(斯皮尔曼相关系数=1)。为确保生成数据基于真实输入而非模型参数知识,我们发布全新 Tempora-0325 数据集(超7000篇2025年3月后发布文档)。涵盖26个主流模型(参数量3–671B),通过算法验证(如引用溯源)与人工评估,全面验证生成评估质量。项目开源包括 YourBench 库、Tempora-0325 数据集、15万+问答对及全部推理日志,支持可复现研究,助力社区按需生成专属评估集。
原文摘要 · Abstract (English)
Evaluating large language models (LLMs) effectively remains a critical bottleneck, as traditional static benchmarks suffer from saturation and contamination, while human evaluations are costly and slow. This hinders timely or domain-specific assessment, crucial for real-world applications. We introduce YourBench, a novel, open-source framework that addresses these limitations by enabling dynamic, automated generation of reliable, up-to-date, and domain-tailored benchmarks cheaply and without manual annotation, directly from user-provided documents. We demonstrate its efficacy by replicating 7 diverse MMLU subsets using minimal source text, achieving this for under 15 USD in total inference costs while perfectly preserving the relative model performance rankings (Spearman Rho = 1) observed on the original benchmark. To ensure that YourBench generates data grounded in provided input instead of relying on posterior parametric knowledge in models, we also introduce Tempora-0325, a novel dataset of over 7K diverse documents, published exclusively after March 2025. Our comprehensive analysis spans 26 SoTA models from 7 major families across varying scales (3-671B parameters) to validate the quality of generated evaluations through rigorous algorithmic checks (e.g., citation grounding) and human assessments. We release the YourBench library, the Tempora-0325 dataset, 150k+ question answer pairs based on Tempora and all evaluation and inference traces to facilitate reproducible research and empower the community to generate bespoke benchmarks on demand, fostering more relevant and trustworthy LLM evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。