arXiv:2504.13261cs.CLcs.AI2025-04被引 1

首个评估大模型中文教学语法能力的多层级基准

CPG-EVAL: A Multi-Tiered Benchmark for Evaluating the Chinese Pedagogical Grammar Competence of Large Language Models

  • 设计五类任务,测试语法识别、辨析与抗干扰能力
  • 小模型在单任务中表现尚可,大模型抗干扰更强但仍有提升空间
  • 为教育场景下大模型应用提供评估依据,适合教研与开发者

大型语言模型(如ChatGPT)的快速发展对外语教学产生了深远影响,但其教学语法能力仍缺乏系统评估。本文提出CPG-EVAL,首个专用于评估大模型在外语教学情境下中文教学语法能力的多层级基准。该基准包含五项任务,涵盖语法识别、细粒度语法区分、类别判别及抗语言干扰能力。实验发现,小规模模型在单一实例任务中表现良好,但在多实例任务和混淆干扰下表现不佳;大规模模型虽具备更强抗干扰能力,但准确率仍有显著提升空间。研究揭示了当前模型在教学对齐上的不足,强调需构建更严谨的评估体系以指导大模型在教育场景中的部署。本工作为教育者、政策制定者及模型开发者提供了实证参考,并为未来提升模型教学适配性与智能化水平奠定基础。

原文摘要 · Abstract (English)

Purpose: The rapid emergence of large language models (LLMs) such as ChatGPT has significantly impacted foreign language education, yet their pedagogical grammar competence remains under-assessed. This paper introduces CPG-EVAL, the first dedicated benchmark specifically designed to evaluate LLMs' knowledge of pedagogical grammar within the context of foreign language instruction. Methodology: The benchmark comprises five tasks designed to assess grammar recognition, fine-grained grammatical distinction, categorical discrimination, and resistance to linguistic interference. Findings: Smaller-scale models can succeed in single language instance tasks, but struggle with multiple instance tasks and interference from confusing instances. Larger-scale models show better resistance to interference but still have significant room for accuracy improvement. The evaluation indicates the need for better instructional alignment and more rigorous benchmarks, to effectively guide the deployment of LLMs in educational contexts. Value: This study offers the first specialized, theory-driven, multi-tiered benchmark framework for systematically evaluating LLMs' pedagogical grammar competence in Chinese language teaching contexts. CPG-EVAL not only provides empirical insights for educators, policymakers, and model developers to better gauge AI's current abilities in educational settings, but also lays the groundwork for future research on improving model alignment, enhancing educational suitability, and ensuring informed decision-making concerning LLM integration in foreign language instruction.

大模型评估教学语法中文教育多层级基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。