arXiv:2507.21419cs.AI2025-07

构建政府领域相关性评估基准,提升大模型在政务场景的能力评测

GovRelBench:A Benchmark for Government Domain Relevance

  • 设计政府领域专用提示与评估工具GovRelBERT
  • 采用软标签训练法提升相关性评分精度
  • 适合政务大模型研发与评测人员使用

当前对大模型在政府领域的评估多聚焦于特定场景下的安全性,而对其核心能力——尤其是领域相关性——的评估仍显不足。为此,我们提出GovRelBench,一个专用于评估大模型在政府领域核心能力的基准。该基准包含政府领域提示数据集及专用评估工具GovRelBERT。在训练GovRelBERT时,引入SoftGovScore方法:基于ModernBERT架构,将硬标签转化为软分数进行训练,使模型能精准计算文本的政府领域相关性得分。本研究旨在完善大模型在政府领域的评估体系,为相关科研与实践提供有效工具。代码与数据集已开源。

原文摘要 · Abstract (English)

Current evaluations of LLMs in the government domain primarily focus on safety considerations in specific scenarios, while the assessment of the models' own core capabilities, particularly domain relevance, remains insufficient. To address this gap, we propose GovRelBench, a benchmark specifically designed for evaluating the core capabilities of LLMs in the government domain. GovRelBench consists of government domain prompts and a dedicated evaluation tool, GovRelBERT. During the training process of GovRelBERT, we introduce the SoftGovScore method: this method trains a model based on the ModernBERT architecture by converting hard labels to soft scores, enabling it to accurately compute the text's government domain relevance score. This work aims to enhance the capability evaluation framework for large models in the government domain, providing an effective tool for relevant research and practice. Our code and dataset are available at https://github.com/pan-xi/GovRelBench.

大模型评测政府领域相关性评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。