arXiv:2506.01419cs.CL2025-06EMNLP被引 13

构建13语言通用语言水平数据集,支持多语言阅读难度自动评估

UniversalCEFR: Enabling Open Multilingual Research on Language Proficiency Assessment

  • 收集50万+文本,统一标注13种语言的CEFR等级
  • 验证语言特征与微调大模型在多语言评估中有效
  • 开源数据格式标准,助力全球研究者公平参与

我们提出UniversalCEFR,一个涵盖13种语言的大规模多维度文本数据集,所有文本均标注了CEFR(欧洲共同语言参考框架)等级。为推动自动化可读性与语言水平评估的研究,该数据集包含505,807条来自教育及学习者资源的文本,已标准化为统一数据格式,支持跨任务、跨语言的一致处理与建模。为验证其价值,我们采用三种建模范式开展基准测试:a) 基于语言特征的分类,b) 微调预训练大语言模型(LLM),c) 使用指令微调后的LLM进行描述引导提示。结果表明,语言特征与微调预训练模型在多语言CEFR评估中表现优异。UniversalCEFR旨在通过标准化数据分布,建立语言水平研究的最佳实践,并向全球研究社区开放数据资源。

原文摘要 · Abstract (English)

We introduce UniversalCEFR, a large-scale multilingual and multidimensional dataset of texts annotated with CEFR (Common European Framework of Reference) levels in 13 languages. To enable open research in automated readability and language proficiency assessment, UniversalCEFR comprises 505,807 CEFR-labeled texts curated from educational and learner-oriented resources, standardized into a unified data format to support consistent processing, analysis, and modelling across tasks and languages. To demonstrate its utility, we conduct benchmarking experiments using three modelling paradigms: a) linguistic feature-based classification, b) fine-tuning pre-trained LLMs, and c) descriptor-based prompting of instruction-tuned LLMs. Our results support using linguistic features and fine-tuning pretrained models in multilingual CEFR level assessment. Overall, UniversalCEFR aims to establish best practices in data distribution for language proficiency research by standardising dataset formats, and promoting their accessibility to the global research community.

语言评估多语言数据集CEFR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。