构建统一动态评测库,支持多领域定制化大模型评估
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
- 整合38个基准数据集,自动分类30.3万道题目
- 实测不同领域下模型表现差异显著,凸显领域感知评估必要性
- 适合关注模型评测公平性与领域适配性的研究人员
随着大语言模型持续发展,构建更新及时、结构清晰的评测基准变得愈发关键。然而现有数据集分散难管,难以满足特定领域或场景的定制化评估需求,尤其在数学、代码等专业领域。本文提出BenchHub,一个动态基准仓库,可聚合并自动分类跨领域评测数据集,集成38个基准中的30.3万道题目。系统支持持续更新与可扩展的数据管理,实现灵活定制的评估方案。通过多类大模型的广泛实验,发现模型在不同领域子集上表现差异明显,强调了领域感知评测的重要性。BenchHub有望促进数据复用、提升模型对比透明度,并识别现有评测中的盲区,为大模型评估研究提供关键基础设施。
原文摘要 · Abstract (English)
As large language models (LLMs) continue to advance, the need for up-to-date and well-organized benchmarks becomes increasingly critical. However, many existing datasets are scattered, difficult to manage, and make it challenging to perform evaluations tailored to specific needs or domains, despite the growing importance of domain-specific models in areas such as math or code. In this paper, we introduce BenchHub, a dynamic benchmark repository that empowers researchers and developers to evaluate LLMs more effectively. BenchHub aggregates and automatically classifies benchmark datasets from diverse domains, integrating 303K questions across 38 benchmarks. It is designed to support continuous updates and scalable data management, enabling flexible and customizable evaluation tailored to various domains or use cases. Through extensive experiments with various LLM families, we demonstrate that model performance varies significantly across domain-specific subsets, emphasizing the importance of domain-aware benchmarking. We believe BenchHub can encourage better dataset reuse, more transparent model comparisons, and easier identification of underrepresented areas in existing benchmarks, offering a critical infrastructure for advancing LLM evaluation research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。