arXiv:2508.07353cs.AIcs.CL2025-08EMNLP被引 3

提出兼顾全面与精简的评测框架,提升专业领域大模型评估效果

Benchmarking for Domain-Specific LLMs: A Case Study on Academia and Beyond

  • 采用全面性与紧凑性双原则,迭代构建评测集
  • 在名校案例中产出大规模高质量学术评测集PolyBench
  • 框架可迁移至医疗、法律等各类专业领域

随着对大语言模型(LLM)领域专用评估需求的增长,大量基准测试被开发。这些工作通常遵循数据规模化的原则,依赖大规模语料或广泛的问答(QA)集以确保覆盖广度。然而,语料和问答集设计对领域专用LLM性能精确率与召回率的影响仍不明确。本文认为,数据规模化并非领域专用评测构建的最优原则。为此,我们提出Comp-Comp框架,基于全面性与紧凑性原则:全面性确保领域语义召回,紧凑性通过减少冗余与噪声提升精确率。为验证方法有效性,我们在一所知名大学开展案例研究,构建出大规模、高质量的学术评测集PolyBench。尽管本研究聚焦学术领域,但Comp-Comp框架具备领域无关性,可广泛适配各类专业场景。代码与数据集详见https://github.com/Anya-RB-Chen/COMP-COMP。

原文摘要 · Abstract (English)

The increasing demand for domain-specific evaluation of large language models (LLMs) has led to the development of numerous benchmarks. These efforts often adhere to the principle of data scaling, relying on large corpora or extensive question-answer (QA) sets to ensure broad coverage. However, the impact of corpus and QA set design on the precision and recall of domain-specific LLM performance remains poorly understood. In this paper, we argue that data scaling is not always the optimal principle for domain-specific benchmark construction. Instead, we introduce Comp-Comp, an iterative benchmarking framework grounded in the principle of comprehensiveness and compactness. Comprehensiveness ensures semantic recall by covering the full breadth of the domain, while compactness improves precision by reducing redundancy and noise. To demonstrate the effectiveness of our approach, we present a case study conducted at a well-renowned university, resulting in the creation of PolyBench, a large-scale, high-quality academic benchmark. Although this study focuses on academia, the Comp-Comp framework is domain-agnostic and readily adaptable to a wide range of specialized fields. The source code and datasets can be accessed at https://github.com/Anya-RB-Chen/COMP-COMP.

大模型评测领域专用基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。