对比六种向量融合方法,发现共享协方差求和最可靠。
Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey

- 采用共享协方差的向量求和法表现最佳
- TSVM对多语言干扰缓解有限,效果不稳定
- 权重缩放与秩压缩比影响显著,需调优
多语言知识编辑(MKE)因语言间编辑相互干扰而困难,即使在单语环境下表现良好的定位-编辑方法也面临挑战。本文研究三个问题:向量融合方法的有效性、任务特异性融合向量(TSVM)减少多语言干扰的能力,以及权重缩放因子和秩压缩比对性能的影响。我们在大规模批量编辑设置下,针对12种语言,在MzsRE基准上,评估了两种主流骨干大模型、两种基础编辑方法与六种融合变体。结果表明,具有共享协方差的向量求和是最可靠的策略,而无共享协方差的简单求和表现较差;TSVM在部分场景下提升性能,但抑制多语言干扰能力有限;性能对权重缩放和秩压缩比敏感,大于默认值的缩放与较低秩常带来更好结果。这些发现厘清了现有向量融合方法在多语言知识编辑中的实际优劣,为未来研究提供指导。
原文摘要 · Abstract (English)
Multilingual knowledge editing (MKE) remains challenging because language-specific edits interfere with one another, even when locate-then-edit methods work well in monolingual settings. This paper focuses on three issues: the effectiveness of vector merging methods for MKE, the extent to which Task Singular Vectors for Merging (TSVM) can reduce multilingual interference, and the influence of the weight scaling factor and rank compression ratio on performance. We evaluate six merging variants with two popular backbone large language models, two base knowledge editing methods, and 12 languages on the MzsRE benchmark under a large-scale batch-editing setting. Our results show that vector summation with shared covariance is the most reliable overall strategy, whereas simple summation without shared covariance performs poorly. TSVM improves performance in some settings, but its ability to mitigate multilingual interference is limited. We also find that performance is sensitive to both weight scale and rank ratio, with larger-than-default scaling and relatively low rank often yielding better results. These findings clarify the practical strengths and limits of current vector merging methods for MKE and provide guidance for future multilingual knowledge editing research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。