arXiv:2502.07938cs.CL2025-02被引 1

让多语言模型学会理解历史卢森堡语,提升跨语言检索效果

Adapting Multilingual Embedding Models to Historical Luxembourgish

  • 用GPT-4o生成2万组历史卢森堡语与德/法/英平行语料
  • 在历史文本上微调后,所有模型跨语言搜索准确率显著提升
  • 适合历史语言处理、低资源语言研究者使用

日益增长的数字化历史文本需要有效的语义搜索。然而,预训练多语言模型因光学字符识别噪声和过时拼写,在处理历史内容时表现不佳。本研究针对历史卢森堡语(LB)这一低资源语言,开展跨语言语义搜索研究。我们收集了不同时期的历史卢森堡语新闻文章,并利用GPT-4o进行句子分割与翻译,每语言对生成20,000组平行语料。此外,构建了一个历史卢森堡语双语语料库评估集,发现现有模型在该任务上表现较差。通过使用我们构建的历史及额外现代平行数据,采用对比学习或知识蒸馏方法对多个多语言嵌入模型进行适配,所有模型的准确率均显著提高。我们公开发布适配后的模型及历史卢森堡语-德/法/英双语语料,以支持后续研究。

原文摘要 · Abstract (English)

The growing volume of digitized historical texts requires effective semantic search using text embeddings. However, pre-trained multilingual models face challenges with historical content due to OCR noise and outdated spellings. This study examines multilingual embeddings for cross-lingual semantic search in historical Luxembourgish (LB), a low-resource language. We collect historical Luxembourgish news articles from various periods and use GPT-4o for sentence segmentation and translation, generating 20,000 parallel training sentences per language pair. Additionally, we create a semantic search (Historical LB Bitext Mining) evaluation set and find that existing models perform poorly on cross-lingual search for historical Luxembourgish. Using our historical and additional modern parallel training data, we adapt several multilingual embedding models through contrastive learning or knowledge distillation and increase accuracy significantly for all models. We release our adapted models and historical Luxembourgish-German/French/English bitexts to support further research.

历史语言多语言低资源语义搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。