让大模型精准生成不同难度的阿拉伯语,助力个性化学习。
Can LLMs Control Readability? A Multi-Dimensional Evaluation Framework for CEFR-Controlled Arabic Generation

- 用结构化提示+词汇约束控制文本难度
- 最高达0.91的相似度与0.99的准确率
- 适合语言教育系统与自适应学习研究
尽管大语言模型能生成流利的阿拉伯语文本,但其对可读性水平的可靠控制仍不明确。本文提出一种多维度评估框架,用于评估基于欧洲共同语言参考框架(CEFR)的阿拉伯语生成效果,检验指令跟随型大模型在自适应语言学习中的可靠性。框架融合受控提示、经验证的Taha-19可读性预测模型、词汇约束验证及句法复杂度分析。结果表明,结构化提示显著提升CEFR一致性。特别是结合词汇约束的CEFR引导提示,在参考语言特征匹配上达到0.91的余弦相似度,且可读性预测一致性达0.99,而无约束提示则控制能力较弱。研究为将可读性感知的阿拉伯语生成融入自适应教育系统提供了实证基础。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) can generate fluent Arabic text, their ability to reliably control readability levels remains unclear. We propose a multi-dimensional evaluation framework for Common European Framework of Reference for Language (CEFR)-controlled Arabic text generation, assessing whether instruction-following LLMs can serve as reliable generators for adaptive language learning. Our framework integrates controlled prompting, automatic readability prediction using a validated Taha-19 model, lexical constraint validation, and syntactic complexity profiling. Results show that structured prompting substantially improves CEFR alignment. In particular, CEFR-guided prompting with lexical constraints achieves the highest conformity to reference linguistic profiles (0.91 cosine similarity) and near-perfect agreement with predicted readability levels (0.99), while unconstrained prompting exhibits weak control. These findings establish an empirical foundation for integrating readability-aware Arabic text generation into adaptive educational systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。