arXiv:2502.19722cs.CLcs.IR2025-02中稿 · TACL

用5个例子让模型学会多语言问答,还能零样本扩展新语言。

Few-Shot Multilingual Open-Domain QA from 5 Examples

  • 用大模型生成高质量多语言数据,仅需少量示例训练。
  • 在多语言问答和跨语言检索上超越现有方法,提升显著。
  • 支持零样本适配新语言,适合资源匮乏语种的部署场景。

多语言开放域问答(MLODQA)在拥有大量语言特定数据时表现良好,但标注成本高限制了其在低资源语言中的应用。本文提出一种少样本学习方法,通过大语言模型生成大规模多语言数据。方法先基于WikiData进行大规模自监督预训练,再使用大模型以少样本提示生成的高质量合成多语言数据进行微调。最终模型FsModQA在多语言问答、跨语言与单语言检索任务上显著优于现有少样本及有监督基线。进一步表明,仅需英语监督数据,结合跨语言提示策略,即可实现新语言的高效零样本适配,为无需大规模标注的多语言问答提供通用解决方案。

原文摘要 · Abstract (English)

Recent approaches to multilingual open-domain question answering (MLODQA) have achieved promising results given abundant language-specific training data. However, the considerable annotation cost limits the application of these methods for underrepresented languages. We introduce a \emph{few-shot learning} approach to synthesise large-scale multilingual data from large language models (LLMs). Our method begins with large-scale self-supervised pre-training using WikiData, followed by training on high-quality synthetic multilingual data generated by prompting LLMs with few-shot supervision. The final model, \textsc{FsModQA}, significantly outperforms existing few-shot and supervised baselines in MLODQA and cross-lingual and monolingual retrieval. We further show our method can be extended for effective zero-shot adaptation to new languages through a \emph{cross-lingual prompting} strategy with only English-supervised data, making it a general and applicable solution for MLODQA tasks without costly large-scale annotation.

少样本学习多语言问答大模型生成零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。