arXiv:2506.03580cs.CL2025-06ACL被引 2

用预训练模型自动生成适合日语学习者的多样例句,提升语言学习体验。

Automatically Suggesting Diverse Example Sentences for L2 Japanese Learners Using Pre-Trained Language Models

  • 结合检索与生成两种方式,利用PLM筛选和创建例句。
  • 检索方法在难度、多样性和自然度上均优于生成方法。
  • 对初学者和进阶者尤其有效,适合语言学习系统开发。

为促进有效语言习得,提供与学习者水平匹配且多样的例句至关重要。本研究探讨使用预训练语言模型(PLMs)为二语日语学习者生成例句。我们采用两种方法:一是在新构建的日语句子语料库中,将PLMs作为质量评分组件进行检索;二是使用零样本学习直接生成句子。通过包含日语学习者、母语者及GPT-4的评审团,从难度、多样性、自然度等多维度评估句子质量。结果表明,各评审者在句子质量评价上存在显著分歧,仅难度评分较为一致。尽管如此,所有评审者更偏好检索方法,尤其适用于初级与高级学习者;而生成方法整体得分较低。实验表明,PLMs有潜力提升句子推荐系统的适应性,从而优化语言学习过程。

原文摘要 · Abstract (English)

Providing example sentences that are diverse and aligned with learners' proficiency levels is essential for fostering effective language acquisition. This study examines the use of Pre-trained Language Models (PLMs) to produce example sentences targeting L2 Japanese learners. We utilize PLMs in two ways: as quality scoring components in a retrieval system that draws from a newly curated corpus of Japanese sentences, and as direct sentence generators using zero-shot learning. We evaluate the quality of sentences by considering multiple aspects such as difficulty, diversity, and naturalness, with a panel of raters consisting of learners of Japanese, native speakers -- and GPT-4. Our findings suggest that there is inherent disagreement among participants on the ratings of sentence qualities, except for difficulty. Despite that, the retrieval approach was preferred by all evaluators, especially for beginner and advanced target proficiency, while the generative approaches received lower scores on average. Even so, our experiments highlight the potential for using PLMs to enhance the adaptability of sentence suggestion systems and therefore improve the language learning journey.

语言学习预训练模型例句生成日语教学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。