首个中文事实性问答基准,评估大模型答短问题的准确性
Chinese SimpleQA: A Chinese Factuality Evaluation for Large Language Models
- 聚焦中文6大主题99个子领域,构建多样化问答数据集
- 严格质量控制确保答案静态可靠,支持自动化评分
- 专为评估中文事实性设计,适合模型开发者优化中文能力
为应对大语言模型快速发展的需求,本文提出首个全面的中文事实性评估基准Chinese SimpleQA,用于评估模型回答简短问题的能力。该基准具有五大特性:中文、多样化、高质量、静态、易评估。首先,在6大主题下覆盖99个多样化子领域;其次,通过全流程质量控制确保问题与答案的高质量,参考答案为静态且不可更改;第三,参照SimpleQA设计,问题和答案均简短,可基于OpenAI API实现便捷评分。基于此基准,我们对现有大模型的事实性能力进行了全面评估,旨在帮助开发者更好理解模型在中文场景下的表现,推动基础模型的发展。
原文摘要 · Abstract (English)
New LLM evaluation benchmarks are important to align with the rapid development of Large Language Models (LLMs). In this work, we present Chinese SimpleQA, the first comprehensive Chinese benchmark to evaluate the factuality ability of language models to answer short questions, and Chinese SimpleQA mainly has five properties (i.e., Chinese, Diverse, High-quality, Static, Easy-to-evaluate). Specifically, first, we focus on the Chinese language over 6 major topics with 99 diverse subtopics. Second, we conduct a comprehensive quality control process to achieve high-quality questions and answers, where the reference answers are static and cannot be changed over time. Third, following SimpleQA, the questions and answers are very short, and the grading process is easy-to-evaluate based on OpenAI API. Based on Chinese SimpleQA, we perform a comprehensive evaluation on the factuality abilities of existing LLMs. Finally, we hope that Chinese SimpleQA could guide the developers to better understand the Chinese factuality abilities of their models and facilitate the growth of foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。