arXiv:2502.11405cs.CL2025-02NAACL被引 12

通过分层融合增强低资源语言推理能力

LayAlign: Enhancing Multilingual Reasoning in Large Language Models via Layer-Wise Adaptive Fusion and Alignment Strategy

  • 融合多语言编码器所有层的表示,实现跨层交互
  • 在多语言推理任务上显著优于现有基线方法
  • 特别适合提升低资源语言的模型表现

尽管大语言模型(LLMs)在多语言语料上预训练,但在低资源语言上表现仍不理想。现有方法通过引入可训练参数连接多语言编码器与LLM,但通常只关注编码器输出层,忽略了其他层的有用信息。本文提出 ame( ame),一个整合多语言编码器所有层表示的框架,并引入 ame 机制,实现LLM与多语言编码器之间的分层交互。在多语言推理任务上的大量实验及表征分析表明,该方法持续优于现有基线。

原文摘要 · Abstract (English)

Despite being pretrained on multilingual corpora, large language models (LLMs) exhibit suboptimal performance on low-resource languages. Recent approaches have leveraged multilingual encoders alongside LLMs by introducing trainable parameters connecting the two models. However, these methods typically focus on the encoder's output, overlooking valuable information from other layers. We propose \aname (\mname), a framework that integrates representations from all encoder layers, coupled with the \attaname mechanism to enable layer-wise interaction between the LLM and the multilingual encoder. Extensive experiments on multilingual reasoning tasks, along with analyses of learned representations, show that our approach consistently outperforms existing baselines.

多语言推理增强分层融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。