用中国职业资格考题评测大模型,发现本地知识更重要。
QualBench: Benchmarking Chinese LLMs with Localized Professional Qualifications for Vertical Domain Evaluation
- 用24个中国职业资格考试题构建多领域测评集
- 中文大模型平均准确率53.98%,部分超越GPT-4o
- 适合评估垂直领域应用能力的开发者与研究者
中文大模型快速发展,亟需垂直领域评估保障应用可靠性。现有基准普遍缺乏领域覆盖,且难以反映中国职场实际需求。本文以职业资格考试为统一评估框架,提出首个面向中文大模型的多领域问答评测集QualBench,涵盖6个垂直领域、超过17,000道题目,源自24项中国职业资格认证,契合国家政策与行业标准。实验显示,中文大模型整体表现优于非中文模型,Qwen2.5在部分任务中超越GPT-4o,凸显本地化知识的重要性。平均准确率为53.98%,反映出当前模型在专业领域仍存在明显短板。此外,研究识别出大模型众包带来的性能下降、数据污染问题,并验证了提示工程与微调的有效性,提示未来可通过多领域RAG与联邦学习实现提升。
原文摘要 · Abstract (English)
The rapid advancement of Chinese LLMs underscores the need for vertical-domain evaluations to ensure reliable applications. However, existing benchmarks often lack domain coverage and provide limited insights into the Chinese working context. Leveraging qualification exams as a unified framework for expertise evaluation, we introduce QualBench, the first multi-domain Chinese QA benchmark dedicated to localized assessment of Chinese LLMs. The dataset includes over 17,000 questions across six vertical domains, drawn from 24 Chinese qualifications to align with national policies and professional standards. Results reveal an interesting pattern of Chinese LLMs consistently surpassing non-Chinese models, with the Qwen2.5 model outperforming the more advanced GPT-4o, emphasizing the value of localized domain knowledge in meeting qualification requirements. The average accuracy of 53.98% reveals the current gaps in domain coverage within model capabilities. Furthermore, we identify performance degradation caused by LLM crowdsourcing, assess data contamination, and illustrate the effectiveness of prompt engineering and model fine-tuning, suggesting opportunities for future improvements through multi-domain RAG and Federated Learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。