arXiv:2506.14046cs.CLcs.AI2025-06被引 4

构建首个对话文本难度评估数据集,助力大模型训练与筛选。

Ace-CEFR -- A Dataset for Automated Evaluation of the Linguistic Difficulty of Conversational Texts for LLM Applications

  • 专家标注英文对话文本的难度等级,形成新数据集
  • 基于该数据集训练的模型精度超人类专家且推理快
  • 适合大模型训练、评估与生产部署场景使用

当前缺乏对短篇对话文本语言难度的自动化评估手段,尤其在大语言模型(LLM)训练与过滤中需求迫切。本文提出 Ace-CEFR,一个由专家标注难度等级的英文对话文本数据集。我们在该数据集上测试了多种模型,包括基于 Transformer 的模型和大语言模型。结果表明,经 Ace-CEFR 训练的模型在文本难度判断上比人类专家更准确,同时具备生产环境所需的低延迟特性。最后,我们公开发布 Ace-CEFR 数据集,供学术研究与技术开发使用。

原文摘要 · Abstract (English)

There is an unmet need to evaluate the language difficulty of short, conversational passages of text, particularly for training and filtering Large Language Models (LLMs). We introduce Ace-CEFR, a dataset of English conversational text passages expert-annotated with their corresponding level of text difficulty. We experiment with several models on Ace-CEFR, including Transformer-based models and LLMs. We show that models trained on Ace-CEFR can measure text difficulty more accurately than human experts and have latency appropriate to production environments. Finally, we release the Ace-CEFR dataset to the public for research and development.

文本难度大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。