arXiv:2601.16390cs.CLcs.AI2026-01被引 4

通过激活调控提升低资源语言模型性能,不改模型权重。

Cross-Lingual Activation Steering for Multilingual Language Models

  • 推理时动态调节神经元激活,实现跨语言能力增强。
  • 分类准确率提升2.3%,生成F1提升3.4%,高资源语言不变差。
  • 适合想提升多语言模型表现但无法重训练的研究者。

大型语言模型具备强大的多语言能力,但主导语言与非主导语言之间仍存在显著性能差距。先前研究认为这是由多语言表征中共享神经元与语言特异性神经元的不平衡所致。我们提出跨语言激活调控(CLAS),一种无需训练的推理阶段干预方法,可选择性地调节神经元激活。在分类与生成基准上评估显示,平均分别获得2.3%(准确率)和3.4%(F1)的提升,同时保持高资源语言性能。我们发现有效迁移源于功能差异而非严格对齐;性能提升与语言聚类分离度增加相关。结果表明,针对性激活调控可在不修改模型权重的情况下释放现有模型的潜在多语言能力。

原文摘要 · Abstract (English)

Large language models exhibit strong multilingual capabilities, yet significant performance gaps persist between dominant and non-dominant languages. Prior work attributes this gap to imbalances between shared and language-specific neurons in multilingual representations. We propose Cross-Lingual Activation Steering (CLAS), a training-free inference-time intervention that selectively modulates neuron activations. We evaluate CLAS on classification and generation benchmarks, achieving average improvements of 2.3% (Acc.) and 3.4% (F1) respectively, while maintaining high-resource language performance. We discover that effective transfer operates through functional divergence rather than strict alignment; performance gains correlate with increased language cluster separation. Our results demonstrate that targeted activation steering can unlock latent multilingual capacity in existing models without modification to model weights.

多语言激活调控推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。