arXiv:2412.17427cs.CL2024-12被引 3

用大模型自动评估儿童故事中词汇的语境信息量,提升识字教育内容生成质量。

Measuring Contextual Informativeness in Child-Directed Text

  • 基于大模型设计自动化评估方法,衡量故事对目标词的语义传达效果。
  • 与人工评分相关性达0.4983,优于最强基线(0.3534)。
  • 方法可推广至成人文本评估,适合教育内容生成研究者使用。

为填补儿童故事词汇教学内容生成中的重要空白,本文研究如何自动评估故事对目标词汇语义的传达能力,该任务对生成教育内容具有重要意义。我们提出「衡量儿童故事中语境信息量」这一任务,给出正式定义并构建了相应数据集。进一步提出一种基于大语言模型(LLM)的自动化方法。实验表明,该方法与人工评价的信息量判断相关性达Spearman 0.4983,显著优于最强基线(0.3534)。额外分析显示,该方法在成人文本语境信息量评估上亦表现优异,且超越所有基线。

原文摘要 · Abstract (English)

To address an important gap in creating children's stories for vocabulary enrichment, we investigate the automatic evaluation of how well stories convey the semantics of target vocabulary words, a task with substantial implications for generating educational content. We motivate this task, which we call measuring contextual informativeness in children's stories, and provide a formal task definition as well as a dataset for the task. We further propose a method for automating the task using a large language model (LLM). Our experiments show that our approach reaches a Spearman correlation of 0.4983 with human judgments of informativeness, while the strongest baseline only obtains a correlation of 0.3534. An additional analysis shows that the LLM-based approach is able to generalize to measuring contextual informativeness in adult-directed text, on which it also outperforms all baselines.

儿童教育语境评估大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。