arXiv:2412.08090cs.CLcs.AI2024-12AAAI被引 3

提升低资源语言的时序推理能力,让多语言大模型更懂时间敏感问题。

Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages

  • 提出CLiTSSA方法,强化跨语言时序语义对齐。
  • 在罗、德、法三语言上表现优于基线,覆盖三种时序任务。
  • 构建mTEMPREASON数据集,支持多级低资源语言研究。

资源丰富语言与低资源语言之间标注资源的巨大差距仍是大语言模型(LLMs)面临的主要障碍。尽管跨语言上下文学习(X-ICL)通过从多语言预训练模型中检索语义对齐示例取得进展,但我们的研究发现,大模型天然偏好同语言语义对齐的跨语言实例,而非直接的跨语言语义对齐,尤其在处理时间敏感查询时表现差异显著。这类查询需要大模型具备良好的时序推理能力,而现有研究主要聚焦英语。为此,本文提出mTEMPREASON时序推理数据集,覆盖不同层级的低资源语言,并设计跨语言时间敏感语义对齐(CLiTSSA)方法,以提升低资源语言中的时序推理能力。我们还构建了包含平行跨语言时序查询及其预期同语言语义相似度评分的数据对。实验证明,相较于现有基线,在罗马尼亚语、德语、法语三种语言上,涵盖三种时序任务及四类主流大模型,CLiTSSA均表现更优,标志着跨语言时序推理资源不平等问题的重要进展。

原文摘要 · Abstract (English)

The unwavering disparity in labeled resources between resource-rich languages and those considered low-resource remains a significant impediment for Large Language Models (LLMs). Recent strides in cross-lingual in-context learning (X-ICL), mainly through semantically aligned examples retrieved from multilingual pre-trained transformers, have shown promise in mitigating this issue. However, our investigation reveals that LLMs intrinsically reward in-language semantically aligned cross-lingual instances over direct cross-lingual semantic alignments, with a pronounced disparity in handling time-sensitive queries in the X-ICL setup. Such queries demand sound temporal reasoning ability from LLMs, yet the advancements have predominantly focused on English. This study aims to bridge this gap by improving temporal reasoning capabilities in low-resource languages. To this end, we introduce mTEMPREASON, a temporal reasoning dataset aimed at the varied degrees of low-resource languages and propose Cross-Lingual Time-Sensitive Semantic Alignment (CLiTSSA), a novel method to improve temporal reasoning in these contexts. To facilitate this, we construct an extension of mTEMPREASON comprising pairs of parallel cross-language temporal queries along with their anticipated in-language semantic similarity scores. Our empirical evidence underscores the superior performance of CLiTSSA compared to established baselines across three languages -- Romanian, German, and French, encompassing three temporal tasks and including a diverse set of four contemporaneous LLMs. This marks a significant step forward in addressing resource disparity in the context of temporal reasoning across languages.

时序推理多语言低资源语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。