arXiv:2604.20666cs.CLcs.AI2026-04中稿 · presentation at En…

专为希腊语和英语设计的检索增强生成嵌入模型,提升跨语言语义对齐效果。

ORPHEAS: A Cross-Lingual Greek-English Embedding Model for Retrieval-Augmented Generation

  • 基于知识图谱构建高质量数据集,针对希腊语复杂形态进行微调。
  • 在多领域语料上训练,实现跨语言检索性能超越现有主流模型。
  • 适合需要高精度双语信息检索的研究者与开发者使用。

在双语希腊语-英语应用中实现高效的检索增强生成,需要能够捕捉领域特定语义关系和跨语言语义对齐的嵌入模型。现有的多语言嵌入模型将表征能力分散于众多语言,限制了其对希腊语的优化,并难以编码希腊语文本中固有的形态复杂性与领域术语结构。本文提出 ORPHEAS,一种专用于希腊语-英语的检索增强生成嵌入模型。ORPHEAS 采用基于知识图谱的微调方法,构建高质量数据集,应用于多样化多领域语料,实现语言无关的语义表示。数值实验在单语和跨语言检索基准上均表明,ORPHEAS 性能优于当前最先进的多语言嵌入模型,证明对形态复杂的语言进行领域专用微调,不会损害跨语言检索能力。

原文摘要 · Abstract (English)

Effective retrieval-augmented generation across bilingual Greek--English applications requires embedding models capable of capturing both domain-specific semantic relationships and cross-lingual semantic alignment. Existing multilingual embedding models distribute their representational capacity across numerous languages, limiting their optimization for Greek and failing to encode the morphological complexity and domain-specific terminological structures inherent in Greek text. In this work, we propose ORPHEAS, a specialized Greek--English embedding model for bilingual retrieval-augmented generation. ORPHEAS is trained with a high quality dataset generated by a knowledge graph-based fine-tuning methodology which is applied to a diverse multi-domain corpus, which enables language-agnostic semantic representations. The numerical experiments across monolingual and cross-lingual retrieval benchmarks reveal that ORPHEAS outperforms state-of-the-art multilingual embedding models, demonstrating that domain-specialized fine-tuning on morphologically complex languages does not compromise cross-lingual retrieval capability.

跨语言嵌入检索增强希腊语处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。