arXiv:2507.09536cs.CL2025-07

用少量数据让定义生成模型支持白俄罗斯语,验证了跨语言适配的可行性。

Adapting Definition Modeling for New Languages: A Case Study on Belarusian

  • 基于4.3万条白俄语定义构建新数据集,支持模型迁移
  • 仅需少量数据即可有效适配新语言,证明方法高效
  • 现有自动评估指标仍不完善,适合词典编纂与低资源语言研究

定义建模旨在上下文中生成新词汇的定义,有助于词典学家记录更多方言和语言。然而,如何利用已有模型支持尚未覆盖的语言仍待探索。本文聚焦于将现有模型适配至白俄罗斯语,提出一个包含43,150条定义的新数据集。实验表明,定义建模系统的适配只需极少数据,但当前自动评估指标在捕捉性能方面仍存在不足。

原文摘要 · Abstract (English)

Definition modeling, the task of generating new definitions for words in context, holds great prospect as a means to assist the work of lexicographers in documenting a broader variety of lects and languages, yet much remains to be done in order to assess how we can leverage pre-existing models for as-of-yet unsupported languages. In this work, we focus on adapting existing models to Belarusian, for which we propose a novel dataset of 43,150 definitions. Our experiments demonstrate that adapting a definition modeling systems requires minimal amounts of data, but that there currently are gaps in what automatic metrics do capture.

定义建模低资源语言词典学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。