arXiv:2410.03182cs.CLcs.AI2024-10被引 9

用大模型生成双语词典例句,助力低资源语言词典编纂

Generating bilingual example sentences with large language models as lexicography assistants

  • 用大模型生成双语例句,结合人类反馈优化质量
  • 高资源语言例句质量好,低资源语言明显下降
  • 可自动评分且适配个人偏好,降低词典制作成本

我们研究了大语言模型在不同资源水平语言(法语-高资源、印尼语-中资源、德顿语-低资源)中生成并评估双语词典例句的表现,以英语为目标语言。基于GDEX标准(典型性、信息量、可理解性)评估生成质量。结果表明,尽管大模型能生成合理例句,但在低资源语言上性能显著下降;同时人类对例句质量的偏好差异大,导致标注者间一致性低。通过上下文学习可有效对齐模型与个体标注者偏好。此外,预训练语言模型可用于自动评分,发现句子困惑度在高资源语言中是典型性和可理解性的良好代理指标。本研究还构建了一个包含600条评分的新型数据集,并揭示了大模型在降低词典编纂成本方面的潜力,尤其适用于低资源语言。

原文摘要 · Abstract (English)

We present a study of LLMs' performance in generating and rating example sentences for bilingual dictionaries across languages with varying resource levels: French (high-resource), Indonesian (mid-resource), and Tetun (low-resource), with English as the target language. We evaluate the quality of LLM-generated examples against the GDEX (Good Dictionary EXample) criteria: typicality, informativeness, and intelligibility. Our findings reveal that while LLMs can generate reasonably good dictionary examples, their performance degrades significantly for lower-resourced languages. We also observe high variability in human preferences for example quality, reflected in low inter-annotator agreement rates. To address this, we demonstrate that in-context learning can successfully align LLMs with individual annotator preferences. Additionally, we explore the use of pre-trained language models for automated rating of examples, finding that sentence perplexity serves as a good proxy for typicality and intelligibility in higher-resourced languages. Our study also contributes a novel dataset of 600 ratings for LLM-generated sentence pairs, and provides insights into the potential of LLMs in reducing the cost of lexicographic work, particularly for low-resource languages.

大模型词典编纂低资源语言自动评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。