首次研究多语言定义建模,发现大模型在无额外训练下表现优于传统方法。
Multilingual Definition Modeling
- 基于四种语言词典数据,微调多语言模型进行单义词定义生成
- 大语言模型在零样本下表现更好,但未实现跨语言协同优势
- 适合对多语言生成与评估感兴趣的研究者
本文首次开展多语言定义建模研究,利用西班牙语、法语、葡萄牙语和德语的单语词典数据,对预训练多语言语言模型在微调后对单义词定义生成的任务表现进行深入实证分析。此外,采用零样本方法测试两种主流聊天型大语言模型在此任务中的多语言能力。结果表明,多语言模型性能可媲美英语,但未能发挥潜在的跨语言协同效应;而大语言模型整体表现更优。对大模型生成定义的全面人工评估揭示了其在零样本和少样本下的能力,也暴露了不足之处。最后,我们发现通过BERTScore衡量的本任务表现与多语言大模型基准任务高度相关,表明该任务是计算资源受限条件下稳定且自然的替代方案。
原文摘要 · Abstract (English)
In this paper, we propose the first multilingual study on definition modeling. We use monolingual dictionary data for four new languages (Spanish, French, Portuguese, and German) and perform an in-depth empirical study to test the performance of pre-trained multilingual language models on definition modeling of monosemic words when finetuned on this data. Furthermore, we use a zero-shot approach to test the multilingual capabilities of two popular chat-based Large Language Models (LLMs) in the task. Results show that multilingual language models can perform on-pair with English but cannot leverage potential cross-lingual synergies, with LLMs generally offering better performance overall. A comprehensive human evaluation of the LLM-generated definition highlights the zero and few-shot capabilities of these models in this new task, also showing their shortcomings. Finally, we show that performance on our task via BERTScore strongly correlates to the performance on multilingual LLM benchmarks, suggesting that our task offers a viable compute-constrained, stable and natural alternative to these.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。