用大模型自动生成适合学习者的简单词义,提升词典编纂效率。
Towards Automated Lexicography: Generating and Evaluating Definitions for Learner's Dictionaries
- 通过大模型迭代简化生成词义,确保语言通俗易懂。
- 在日语学习词典数据集上,生成结果与人工评价高度一致。
- 构建新评估体系,让大模型担任评判员,提升自动化水平。
本文研究词典定义生成(DDG),即为词条生成非上下文相关的定义。词典定义是学习词义的重要资源,但人工编写成本高,因此亟需自动化。本文聚焦学习者词典定义生成(LDDG),要求定义使用简单词汇。首先,提出一种基于大模型判官的新评估方法,结合自建的日语学习词典数据集(与专业词典学家合作构建),验证结果显示该方法与人工标注具有较好一致性。其次,提出一种基于大模型迭代简化的LDDG方法,实验表明生成的定义在评估标准上得分高,同时保持了词汇的简洁性。
原文摘要 · Abstract (English)
We study dictionary definition generation (DDG), i.e., the generation of non-contextualized definitions for given headwords. Dictionary definitions are an essential resource for learning word senses, but manually creating them is costly, which motivates us to automate the process. Specifically, we address learner's dictionary definition generation (LDDG), where definitions should consist of simple words. First, we introduce a reliable evaluation approach for DDG, based on our new evaluation criteria and powered by an LLM-as-a-judge. To provide reference definitions for the evaluation, we also construct a Japanese dataset in collaboration with a professional lexicographer. Validation results demonstrate that our evaluation approach agrees reasonably well with human annotators. Second, we propose an LDDG approach via iterative simplification with an LLM. Experimental results indicate that definitions generated by our approach achieve high scores on our criteria while maintaining lexical simplicity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。