arXiv:2410.15144cs.CL2024-10综述被引 1

综述神经网络如何用可比数据自动找双语词对,促进建立更精准的双语词典。

A survey of neural-network-based methods utilising comparable data for finding translation equivalents

  • 基于可比语料的神经网络方法自动挖掘双语词对
  • 从词典学角度分析现有方法并提出改进方向
  • 适合对双语词典构建和NLP交叉研究感兴趣的学者

在众多自然语言处理(NLP)应用中,构建双语词典组件的重要性毋庸置疑。然而,词典编制过程耗时费力,且需结合NLP与词典学两领域知识,而前者常忽略后者。本文系统梳理了主流NLP方法中用于自动获取关键词典成分——翻译等价词的方法,重点聚焦基于神经网络、利用可比语料的技术。我们从词典学视角分析这些方法的观点,并识别出融合词典学理念、可在多种需翻译等价词的应用中进一步拓展的方法。本综述旨在推动NLP与词典学领域的联动,使NLP能汲取词典学洞见,同时为基于可比数据的神经网络方法研究提供启发性参考。

原文摘要 · Abstract (English)

The importance of inducing bilingual dictionary components in many natural language processing (NLP) applications is indisputable. However, the dictionary compilation process requires extensive work and combines two disciplines, NLP and lexicography, while the former often omits the latter. In this paper, we present the most common approaches from NLP that endeavour to automatically induce one of the essential dictionary components, translation equivalents and focus on the neural-network-based methods using comparable data. We analyse them from a lexicographic perspective since their viewpoints are crucial for improving the described methods. Moreover, we identify the methods that integrate these viewpoints and can be further exploited in various applications that require them. This survey encourages a connection between the NLP and lexicography fields as the NLP field can benefit from lexicographic insights, and it serves as a helping and inspiring material for further research in the context of neural-network-based methods utilising comparable data.

双语词典神经网络可比语料词典学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。