arXiv:2410.07239cs.CL2024-10EMNLP被引 3

提出新方法评估跨语言词汇对齐,发现现有模型仍有提升空间。

Locally Measuring Cross-lingual Lexical Alignment: A Domain and Word Level Perspective

  • 从词汇和领域层面分析跨语言对齐,引入语境嵌入新指标。
  • 在16种语言上验证,发现主流模型对齐效果仍有显著提升空间。
  • 适用于多语言研究、翻译与跨文化语义分析的学者。

以往自然语言处理中对词汇表示空间的对齐研究主要关注整体语言空间的对齐,而认知科学则更关注局部视角:翻译对应词是否真正意义一致,以及文化与地域差异带来的语义变化。随着技术进步和数据增多,这一长期问题可采用更数据驱动的方式研究。然而,缺乏有效的评估方法制约了进展。本文填补该空白,提出一种结合合成验证与新型自然主义验证(基于亲属称谓领域的词汇空缺)的方法,并引入基于上下文嵌入的新度量指标。研究覆盖16种多样语言,结果表明使用更先进的语言模型能显著改善对齐效果,为实现更精确、细致的跨语言词汇对齐方法与评估提供了新路径。

原文摘要 · Abstract (English)

NLP research on aligning lexical representation spaces to one another has so far focused on aligning language spaces in their entirety. However, cognitive science has long focused on a local perspective, investigating whether translation equivalents truly share the same meaning or the extent that cultural and regional influences result in meaning variations. With recent technological advances and the increasing amounts of available data, the longstanding question of cross-lingual lexical alignment can now be approached in a more data-driven manner. However, developing metrics for the task requires some methodology for comparing metric efficacy. We address this gap and present a methodology for analyzing both synthetic validations and a novel naturalistic validation using lexical gaps in the kinship domain. We further propose new metrics, hitherto unexplored on this task, based on contextualized embeddings. Our analysis spans 16 diverse languages, demonstrating that there is substantial room for improvement with the use of newer language models. Our research paves the way for more accurate and nuanced cross-lingual lexical alignment methodologies and evaluation.

跨语言对齐词汇语义上下文嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。