arXiv:2604.18106cs.CL2026-04ACL

用三源动态融合提升低资源语言模型性能,无需标注数据

Efficient Low-Resource Language Adaptation via Multi-Source Dynamic Logit Fusion

论文配图:Efficient Low-Resource Language Adaptation via Multi-Source Dynamic Logit Fusion
图 1 · 摘自论文原文
  • 测试时融合小模型、高资源任务模型和大模型的预测结果
  • 在8种低资源语言上超越单模型基线和代理微调方法
  • 强调小模型在低资源语言上的主导作用,挑战大模型主导思维

将大语言模型(LLMs)适配到低资源语言(LRLs)受限于任务数据稀缺与计算资源不足。尽管代理微调(Proxy Tuning)提供了一种在输出层融合知识的策略,但在低资源场景中常因大模型对低资源语言能力较弱而失效,反而掩盖了小型专用模型的知识。为此,我们提出TriMix,一种测试时的动态逻辑融合框架,通过整合三个来源的能力:持续预训练的小模型在低资源语言上的表现、高资源语言指令微调模型的任务能力,以及大模型的规模优势。该方法仅需小模型的持续预训练,无需低资源语言任务标注,具备数据与计算高效性。在四个模型家族和八种低资源语言上的实验表明,TriMix始终优于单模型基线和代理微调。分析显示,优先采纳小模型的逻辑输出是成功关键,挑战了以大模型为主导的普遍假设。

原文摘要 · Abstract (English)

Adapting large language models (LLMs) to low-resource languages (LRLs) is constrained by the scarcity of task data and computational resources. Although Proxy Tuning offers a logit-level strategy for introducing scaling effects, it often fails in LRL settings because the large model's weak LRL competence might overwhelm the knowledge of specialized smaller models. We thus propose TriMix, a test-time logit fusion framework that dynamically balances capabilities from three different sources: LRL competence from a continually pretrained small model, task competence from high-resource language instruction tuning, and the scaling benefits of large models. It is data- and compute-efficient, requiring no LRL task annotations, and only continual pretraining on a small model. Experiments across four model families and eight LRLs show that TriMix consistently outperforms single-model baselines and Proxy Tuning. Our analysis reveals that prioritizing the small LRL-specialized model's logits is crucial for success, challenging the prevalent large-model-dominant assumption.

低资源语言模型融合动态推理小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。