arXiv:2603.25489cs.CL2026-03被引 1

利用大模型翻译不对称性,提升罗曼什语低资源方言的机器翻译效果。

Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties

  • 利用大模型从罗曼什语向德语翻译生成合成数据,反向增强低资源方言
  • 在最弱资源方言上比Gemini 3 Pro高出23 BLEU,首次实现流畅方言翻译
  • 适用于低资源语言数据匮乏场景,尤其适合多方言语言的翻译研究

当前低资源机器翻译常依赖大模型基于高资源语言文本生成合成数据。本文针对罗曼什语的6种方言重新审视该方法:大模型在将方言翻译为罗曼什时易混淆,但将罗曼什翻译为德语等高资源语言时表现优异。由于这种翻译不对称性,数据增强方向至关重要。实验发现,与近期策略相反,生成向高资源语言的合成翻译更优,仅此方法使德国-罗曼什翻译性能超越Gemini 3 Pro基准线(最低资源方言+23 BLEU)。人工评估确认,本方法是首个能生成各罗曼什方言流畅翻译的模型。

原文摘要 · Abstract (English)

Recent strategies for low-resource machine translation rely on LLMs to generate synthetic data based on text in higher-resource languages. We revisit this idea for Romansh, a language with 6 distinct varieties. LLMs tend to confuse these varieties when translating into Romansh, but they are quite good at translating out of Romansh into a high-resource language such as German. Due to this asymmetry, the direction of data augmentation is a crucial choice. We find that contrary to recent strategies, creating synthetic translations into the higher-resource language is the superior approach, and only this approach allows us to surpass a Gemini 3 Pro baseline on German-Romansh translation (+23 BLEU over Gemini in the lowest-resource variety). A human evaluation confirms that our experiments yield the first model that generates fluent translations in the individual Romansh varieties.

机器翻译低资源语言大模型应用多方言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。