构建多语言地缘争端数据集,评测跨语言RAG的鲁棒性
Multilingual Retrieval Augmented Generation for Culturally-Sensitive Tasks: A Benchmark for Cross-lingual Robustness
- 构建49语言地缘争端数据集,评估多语言检索增强生成效果
- 跨语言检索显著提升响应一致性,降低地缘政治偏见
- 低资源语言查询时引用分布更广,揭示信息获取不平等
检索增强生成(RAG)有助于缓解大语言模型的幻觉问题,但也会引入检索文档中的偏见。在多语言且文化敏感的任务中(如领土争端),这种偏见可能被放大。为此,我们构建了BordIRLines数据集,包含49种语言的领土争端案例及其对应的维基百科检索文档。通过形式化多种多语言检索模式,我们在多个大语言模型上评估了RAG的跨语言鲁棒性。实验表明,融入多元语言视角能提升系统鲁棒性;使用多语言检索文档相比纯同语种检索,显著提高了响应一致性并降低了地缘政治偏见。我们还分析了RAG响应对文档的利用方式,发现以低资源语言查询时,回答引用的语言分布更为广泛。进一步分析涵盖从检索到文档内容的跨语言RAG全流程。相关基准与代码已开源,网址:https://huggingface.co/datasets/borderlines/bordirlines。
原文摘要 · Abstract (English)
The paradigm of retrieval-augmented generated (RAG) helps mitigate hallucinations of large language models (LLMs). However, RAG also introduces biases contained within the retrieved documents. These biases can be amplified in scenarios which are multilingual and culturally-sensitive, such as territorial disputes. We thus introduce BordIRLines, a dataset of territorial disputes paired with retrieved Wikipedia documents, across 49 languages. We evaluate the cross-lingual robustness of this RAG setting by formalizing several modes for multilingual retrieval. Our experiments on several LLMs show that incorporating perspectives from diverse languages can in fact improve robustness; retrieving multilingual documents best improves response consistency and decreases geopolitical bias over RAG with purely in-language documents. We also consider how RAG responses utilize presented documents, finding a much wider variance in the linguistic distribution of response citations, when querying in low-resource languages. Our further analyses investigate the various aspects of a cross-lingual RAG pipeline, from retrieval to document contents. We release our benchmark and code to support continued research towards equitable information access across languages at https://huggingface.co/datasets/borderlines/bordirlines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。