arXiv:2602.01162cs.CL2026-02

用语言类型学优化大模型翻译低资源语言,不需训练数据

Typologically-Informed Candidate Reranking for LLM-based Translation into Low-Resource Languages

  • 基于16个语言类型维度构建结构化语言画像,量化与英语的差异
  • 对341句测试句干预精准率达86.26%,与语言类型距离强相关
  • 无需平行语料,适配任意可生成多候选的大模型,适合低资源语言

主流大语言模型因主要在高资源语言上训练,对类型学差异大的低资源语言翻译时存在系统性偏差。本文提出无需并行数据或模型重训练的框架,包含两大组件:通用元语言框架(UMF)以16个类型学维度刻画语言特征并加权打分;计算引擎在生成阶段进行语言消歧,在选择阶段评估类型符合度。跨九种语言对的评估显示干预率与英语类型距离高度相关。在341个含不同形态与句法现象的英文句子上,保守处理语言的干预精准率达48.16%,形态密集语言为28.15%,结构可画像语言达86.26%。该框架无需平行数据,兼容任何能生成多候选输出的LLM,可直接部署于低资源语言场景。

原文摘要 · Abstract (English)

Large language models trained predominantly on high-resource languages exhibit systematic biases toward dominant typological patterns, leading to structural non-conformance when translating into typologically divergent low-resource languages. We present a framework that leverages linguistic typology to improve translation quality without parallel training data or model retraining. The framework consists of two components: the Universal Metalinguistic Framework (UMF), which represents languages as structured profiles across 16 typological dimensions with divergence-weighted scoring, and the Computational Engine, which operates through linguistic disambiguation during generation and typological compliance scoring during selection. Evaluation across nine language pairs demonstrates intervention rates strongly correlating with typological distance from English. In experiments on 341 English sentences each having different morphological and syntactic phenomena, the framework shows an intervention precision of 48.16% for conservatively treated languages, 28.15% for morphologically dense languages, and 86.26% for structurally profiled languages. The framework requires no parallel training data and operates with any LLM capable of producing multiple candidate outputs, enabling practical deployment for under-resourced languages.

机器翻译低资源语言语言类型学大模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。