arXiv:2501.04858cs.CL2025-01被引 3

针对波斯语构建更精准的检索增强生成系统,提升低资源语言AI应用能力。

Advancing Retrieval-Augmented Generation for Persian: Development of Language Models, Comprehensive Benchmarks, and Best Practices for Optimization

  • 专为波斯语设计语言模型MatinaRoberta与MatinaSRoberta,适配复杂语法与形态。
  • 在731亿词元语料上训练,大模型(如Llama-3.1 70B)生成准确率最高。
  • 优化检索策略可显著提升法律与专业文本的问答表现,适合低资源语言研究者。

本文研究了在低资源语言中构建检索增强生成(RAG)系统的挑战,聚焦波斯语复杂的形态学与多变句法。通过引入专用于波斯语的模型——掩码语言模型MatinaRoberta和微调版Sentence-BERT MatinaSRoberta,以及一个全面的评估框架,提升了检索与生成精度。模型在包含73.11亿波斯语词元的多样化语料上进行预训练,并使用定制损失函数进行微调。三个数据集(通用知识PQuad、科学文献、组织报告)用于评估:结果显示,MatinaSRoberta在所有数据集上均优于现有嵌入方法,显著提高上下文相关性与检索准确率。通过温度调节、分块大小调整与文档摘要索引等优化手段进一步改进RAG性能。较大模型(如Llama-3.1 70B)始终表现出最高生成准确率,而小型模型在领域特定和正式语境下表现受限。研究证实,通过定制化嵌入与检索-生成配置,可在波斯语等低资源语言中有效构建高质量RAG系统,推动搜索引擎与法律文档分析等应用发展。

原文摘要 · Abstract (English)

This paper examines the specific obstacles of constructing Retrieval-Augmented Generation(RAG) systems in low-resource languages, with a focus on Persian's complicated morphology and versatile syntax. The research aims to improve retrieval and generation accuracy by introducing Persian-specific models, namely MatinaRoberta(a masked language model) and MatinaSRoberta(a fine-tuned Sentence-BERT), along with a comprehensive benchmarking framework. Three datasets-general knowledge(PQuad), scientifically specialized texts, and organizational reports, were used to assess these models after they were trained on a varied corpus of 73.11 billion Persian tokens. The methodology involved extensive pretraining, fine-tuning with tailored loss functions, and systematic evaluations using both traditional metrics and the Retrieval-Augmented Generation Assessment framework. The results show that MatinaSRoberta outperformed previous embeddings, achieving superior contextual relevance and retrieval accuracy across datasets. Temperature tweaking, chunk size modifications, and document summary indexing were explored to enhance RAG setups. Larger models like Llama-3.1 (70B) consistently demonstrated the highest generation accuracy, while smaller models faced challenges with domain-specific and formal contexts. The findings underscore the potential for developing RAG systems in Persian through customized embeddings and retrieval-generation settings and highlight the enhancement of NLP applications such as search engines and legal document analysis in low-resource languages.

检索增强波斯语低资源语言生成优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。