arXiv:2608.05138eess.AScs.AI2026-08被引 1

为现代希腊语构建首个大规模RAG基准,适配检索生成系统。

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

  • 从专业领域数据中挖掘语料,用合成监督训练希腊语检索模型。
  • 微调后模型nDCG@10提升至0.835,问答正确率从29.4%增至66.9%。
  • 发布首个希腊语RAG基准HERA,支持未来多领域研究。

现代希腊语缺失于NVIDIA Nemotron检索模型及主流多语言检索基准,但对法律、能源、金融和医疗等领域的检索增强生成(RAG)至关重要。本文提出端到端适配Nemotron检索栈的完整方案,涵盖语料挖掘、合成监督、检索模型训练、重排序器适配、阅读器微调,并构建首个希腊语RAG基准HERA。研究表明,无需参数的BM25基线在专业希腊语语料上优于多个现成多语言密集检索模型。在65,773个希腊语检索对上微调后,Nemotron 1B嵌入器将nDCG@10从0.362提升至0.835,显著优于原始版本。所学语言能力可迁移到通用领域,但优势仍具领域依赖性。进一步适配交叉编码器重排序器,在各专业领域均实现一致提升。最后,使用LoRA微调Nemotron 30B-A3B混合专家阅读器,使人工评估答案正确率从29.4%升至66.9%,同时大幅改善忠实度与引用质量。本文还发布适应后的模型与基准HERA,以支持未来希腊语RAG研究。

原文摘要 · Abstract (English)

Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation of the Nemotron retrieval stack for Modern Greek, including corpus mining, synthetic supervision, retrieval model training, reranker adaptation, reader fine-tuning, and a new benchmark called HERA. Our study shows that a parameter-free BM25 baseline outperforms several off-the-shelf multilingual dense retrieval models on specialist Greek corpora. After fine-tuning on 65,773 Greek retrieval pairs, a Nemotron 1B embedder improves nDCG@10 from 0.362 to 0.835 and substantially outperforms its unadapted counterpart. The learned language competence transfers to general-domain Greek, although the advantage over BM25 remains domain-dependent. We further adapt a cross-encoder reranker and demonstrate consistent improvements across specialist domains. Finally, we LoRA-tune a Nemotron 30B-A3B mixture-of-experts reader for grounded generation, increasing judged answer correctness from 29.4% to 66.9% while significantly improving faithfulness and citation quality. We also introduce HERA, the first large-scale Greek benchmark for retrieval-augmented generation, and release our adapted models and benchmark to support future research on Greek-language RAG systems.

希腊语RAG检索增强多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。