arXiv:2501.09158cs.CL2025-01被引 6

用新方法评估大模型简化教育文本的效果,提升可读性同时保关键信息。

Evaluating GenAI for Simplifying Texts for Education: Improving Accuracy and Consistency for Enhanced Readability

  • 设计多智能体架构与统一评测指标,系统评估简化效果。
  • 模型在四至八年级文本简化中准确率差异显著,关键词保留率波动大。
  • 适合教育科技研发者与个性化学习工具开发者参考。

生成式人工智能在支持个性化学习方面潜力巨大。教师需要高效工具将教育文本调整至不同阅读水平,同时保留关键内容。大语言模型(LLMs)有潜力满足此需求,但现有方法存在多项不足。本研究提出一种通用评估方法与指标,系统评估了多种LLM、提示技术及新型多智能体架构在简化60篇信息类阅读材料中的表现,目标是将原文从12年级水平降至8、6、4年级水平。计算了各模型与提示技术在目标年级准确度、词汇量变化百分比、关键词与关键短语保持程度(语义相似度)上的表现。单样本t检验与多元回归分析显示,不同模型和提示技术在四项指标上均有显著差异。无论是模型还是提示方法,在降低至4年级水平时,年级准确性与关键词一致性均表现不一。结果表明,LLM在自动化文本简化中具有前景,但当前模型与提示方法难以在各项评估标准间取得理想平衡,并验证了一种可推广的未来系统评估方法。

原文摘要 · Abstract (English)

Generative artificial intelligence (GenAI) holds great promise as a tool to support personalized learning. Teachers need tools to efficiently and effectively enhance content readability of educational texts so that they are matched to individual students reading levels, while retaining key details. Large Language Models (LLMs) show potential to fill this need, but previous research notes multiple shortcomings in current approaches. In this study, we introduced a generalized approach and metrics for the systematic evaluation of the accuracy and consistency in which LLMs, prompting techniques, and a novel multi-agent architecture to simplify sixty informational reading passages, reducing each from the twelfth grade level down to the eighth, sixth, and fourth grade levels. We calculated the degree to which each LLM and prompting technique accurately achieved the targeted grade level for each passage, percentage change in word count, and consistency in maintaining keywords and key phrases (semantic similarity). One-sample t-tests and multiple regression models revealed significant differences in the best performing LLM and prompt technique for each of the four metrics. Both LLMs and prompting techniques demonstrated variable utility in grade level accuracy and consistency of keywords and key phrases when attempting to level content down to the fourth grade reading level. These results demonstrate the promise of the application of LLMs for efficient and precise automated text simplification, the shortcomings of current models and prompting methods in attaining an ideal balance across various evaluation criteria, and a generalizable method to evaluate future systems.

文本简化教育AI大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。