通过可调节融合提升多语言推理模型性能
Enhancing Multilingual Reasoning via Steerable Model Merging

- 引入门控交叉注意力动态调整源模型贡献
- 在21种语言的4个基准上超越多个强基线
- 适合需要灵活适配不同输入的多语言任务
模型融合能有效结合多语言模型与推理模型的能力,通过对齐不同模型的特征空间,在多语言推理任务中实现良好泛化。然而,合并后的单一模型常因源模型间冲突导致性能不佳。现有‘一刀切’的融合策略未能考虑不同输入对模型需求的差异。为此,我们提出可调节模型融合(ST-Merge)框架,通过门控交叉注意力机制自适应地加权或过滤两个源模型的输出。大量实验表明,ST-Merge在跨21种语言的四个多语言推理基准上持续优于多个强基线。
原文摘要 · Abstract (English)
Model merging is an effective technique for composing the capabilities of a multilingual model and a reasoning model. It has achieved promising generalization in multilingual reasoning tasks by aligning feature spaces of different models. However, the merged single model often fails to address the conflicts between source models, leading to suboptimal performance. In other words, the one-size-fits-all merging strategy may not align with the characteristics of different inputs which may require prioritizing certain models over others. To this end, we propose a Steerable Model Merging (ST-Merge) framework to modulate the contribution of each source model. To realize this idea, we introduce a gated cross-attention mechanism to weight or filter the two attended source models in an adaptive manner. Extensive experiments demonstrate that ST-Merge consistently outperforms multiple strong baselines on four multilingual reasoning benchmarks across 21 different languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。