arXiv:2508.01918cs.CLcs.AI2025-08被引 1

首个开源旁遮普语生成模型,用量子启发检索提升低资源语言性能。

Quantum-RAG and PunGPT2: Advancing Low-Resource Language Generation and Retrieval for the Punjabi Language

  • 构建旁遮普语专用分词器与35GB语料库,训练出PunGPT2生成模型。
  • 量子检索框架使召回率提升7.4%,在新评测集上超越多语言基线。
  • 开源全部代码、数据与模型,助力低资源语言研究与应用。

尽管大型语言模型发展迅速,低资源语言仍被排除在自然语言处理之外,限制了数百万用户的数字接入。我们提出PunGPT2,首个完全开源的旁遮普语生成模型套件,基于35GB语料库训练,涵盖文学、宗教文本、新闻及社交话语等。PunGPT2通过针对古鲁姆基和沙赫穆基文字的优化分词器,捕捉旁遮普语的句法与形态丰富性。我们引入Pun-RAG,一个集成PunGPT2与FAISS检索器的检索增强框架,以及使用QLoRA微调的Pun-Instruct变体,实现强大的零样本摘要、翻译与问答能力。关键创新为Quantum-RAG,融合稀疏、密集与量子核嵌入,实现高效、上下文感知的检索,内存开销低,是首个在低资源大模型中实现的实用量子启发检索。模型在FLORES-200、IndicGenBench及新设的PunjabiEval评测集上优于多语言基线(mBERT、mT5、MuRIL、BLOOM)。Quantum-RAG在PunjabiEval上相较FAISS提升+7.4 Recall@10,较mT5提升+3.5 BLEU。我们公开发布所有训练脚本、超参数、评估流程、35GB旁遮普语语料库、PunjabiEval基准及全部模型权重,确立旁遮普语生成与检索的新基准。

原文摘要 · Abstract (English)

Despite rapid advances in large language models (LLMs), low-resource languages remain excluded from NLP, limiting digital access for millions. We present PunGPT2, the first fully open-source Punjabi generative model suite, trained on a 35GB corpus covering literature, religious texts, news, social discourse, etc. PunGPT2 captures Punjabi's syntactic and morphological richness through a tokenizer optimized for Gurmukhi and Shahmukhi scripts. We introduce Pun-RAG, a retrieval-augmented framework integrating PunGPT2 with a FAISS retriever over a curated Punjabi knowledge base, and Pun-Instruct, an instruction-tuned variant using QLoRA for robust zero-shot summarization, translation, and question answering. Our key innovation, Quantum-RAG, fuses sparse, dense, and quantum kernel embeddings for efficient, context-aware retrieval with low memory overhead, marking the first practical quantum-inspired retrieval in a low-resource LLM. Our models outperform multilingual baselines (mBERT, mT5, MuRIL, BLOOM) on FLORES-200, IndicGenBench, and a new PunjabiEval suite. Quantum-RAG yields +7.4 Recall@10 over FAISS and +3.5 BLEU over mT5 on PunjabiEval. We publicly release all training scripts, hyperparameters, evaluation pipelines, the 35GB Punjabi corpus, the PunjabiEval benchmark, and all model weights, establishing new state-of-the-art results for Punjabi language generation and retrieval.

低资源语言生成模型检索增强量子启发

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。