arXiv:2410.21970cs.CL2024-10被引 8

发现多语言检索增强生成中存在语言不平等,英语占优,低资源语言易出错。

Not All Languages are Equal: Insights into Multilingual Retrieval-Augmented Generation

  • 构建八语言平行语料库Futurepedia,评估多语言RAG模型表现
  • 英语在知识提取与选择中占据优势,印欧语系更易直接引用原文
  • 建议增加非英语数据、调整文档权重以减少语言偏见

RALMs通过引入外部文本扩展知识范围,但全球知识的多语言特性要求其处理多种语言,这一问题研究不足。本文提出Futurepedia基准,包含八种代表性语言的平行文本,评估六种多语言RALMs。实验揭示语言不平等:1)高资源语言在单语知识提取中表现更优;2)印欧语系使模型更倾向于直接引用原文,降低跨语言表达难度;3)英语因模型选择偏见而更具影响力。据此提出改进建议:单语知识提取需关注低资源语言翻译带来的误差传递;跨语言知识迁移应鼓励模型在不同语言文档中直接作答;多语言知识选择应增加非英语文档并调整英语文档权重。通过全面实验,揭示多语言RALMs的复杂性,为未来研究提供关键洞察。

原文摘要 · Abstract (English)

RALMs (Retrieval-Augmented Language Models) broaden their knowledge scope by incorporating external textual resources. However, the multilingual nature of global knowledge necessitates RALMs to handle diverse languages, a topic that has received limited research focus. In this work, we propose \textit{Futurepedia}, a carefully crafted benchmark containing parallel texts across eight representative languages. We evaluate six multilingual RALMs using our benchmark to explore the challenges of multilingual RALMs. Experimental results reveal linguistic inequalities: 1) high-resource languages stand out in Monolingual Knowledge Extraction; 2) Indo-European languages lead RALMs to provide answers directly from documents, alleviating the challenge of expressing answers across languages; 3) English benefits from RALMs' selection bias and speaks louder in multilingual knowledge selection. Based on these findings, we offer advice for improving multilingual Retrieval Augmented Generation. For monolingual knowledge extraction, careful attention must be paid to cascading errors from translating low-resource languages into high-resource ones. In cross-lingual knowledge transfer, encouraging RALMs to provide answers within documents in different languages can improve transfer performance. For multilingual knowledge selection, incorporating more non-English documents and repositioning English documents can help mitigate RALMs' selection bias. Through comprehensive experiments, we underscore the complexities inherent in multilingual RALMs and offer valuable insights for future research.

多语言RAG语言偏见知识提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。