arXiv:2512.06483cs.CLcs.AI2025-12中稿 · 3rd International …被引 1

用大模型自动划分德语水平,准确率超以往方法。

Classifying German Language Proficiency Levels Using Large Language Models

  • 融合真实与合成数据构建德语水平标注集
  • 微调LLaMA-3-8B模型在CEFR分类上表现最优
  • 无需修改模型即可通过内部状态判别语言水平

语言能力评估对教育至关重要,有助于实现个性化教学。本文研究使用大型语言模型(LLMs)自动将德语文本按欧洲共同语言参考框架(CEFR)划分为不同语言水平。为支持稳健训练与评估,我们结合多个现有CEFR标注语料库与合成数据,构建了多样化数据集。进一步评估了提示工程、微调LLaMA-3-8B-Instruct模型以及基于探针的内部神经状态分析方法。结果表明,该方法在性能上持续优于先前方法,凸显了大模型在可靠、可扩展的CEFR分类中的潜力。

原文摘要 · Abstract (English)

Assessing language proficiency is essential for education, as it enables instruction tailored to learners needs. This paper investigates the use of Large Language Models (LLMs) for automatically classifying German texts according to the Common European Framework of Reference for Languages (CEFR) into different proficiency levels. To support robust training and evaluation, we construct a diverse dataset by combining multiple existing CEFR-annotated corpora with synthetic data. We then evaluate prompt-engineering strategies, fine-tuning of a LLaMA-3-8B-Instruct model and a probing-based approach that utilizes the internal neural state of the LLM for classification. Our results show a consistent performance improvement over prior methods, highlighting the potential of LLMs for reliable and scalable CEFR classification.

语言评估大模型应用德语学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。