arXiv:2510.26024cs.CLcs.AI2025-10被引 4

提出新方法平衡多语言模型的知识迁移与文化特异性

Rethinking Cross-lingual Alignment: Balancing Transfer and Cultural Erasure in Multilingual LLMs

  • 通过分层激活控制,分离通用知识与文化特异性表达
  • 实验证明现有方法在6种语言上均牺牲文化差异性换取知识迁移
  • 适合关注多语言模型公平性与文化适配的研究者使用

跨语言对齐(CLA)旨在统一多语言表征,使大语言模型能无缝实现跨语言知识迁移。然而,我们提出假设:追求表征趋同可能无意导致‘文化消解’——即丧失基于查询语言提供文化相关回应的能力。本文构建了综合评估框架‘迁移-定位平面’,量化知识迁移与文化消解程度。在六种语言上的实验表明,现有CLA方法虽提升事实知识转移,却显著损害文化定位能力。分析模型内部表征发现,通用知识迁移与文化特异性知识在不同模型层具有最优可操控性。基于此,我们提出推理时的‘外科操纵’(Surgical Steering)方法,通过针对性地调节不同层的激活值,有效平衡两者矛盾,突破现有对齐技术局限。

原文摘要 · Abstract (English)

Cross-lingual alignment (CLA) aims to align multilingual representations, enabling Large Language Models (LLMs) to seamlessly transfer knowledge across languages. While intuitive, we hypothesize, this pursuit of representational convergence can inadvertently cause "cultural erasure", the functional loss of providing culturally-situated responses that should diverge based on the query language. In this work, we systematically analyze this trade-off by introducing a holistic evaluation framework, the transfer-localization plane, which quantifies both desirable knowledge transfer and undesirable cultural erasure. Using this framework, we re-evaluate recent CLA approaches and find that they consistently improve factual transfer at the direct cost of cultural localization across all six languages studied. Our investigation into the internal representations of these models reveals a key insight: universal factual transfer and culturally-specific knowledge are optimally steerable at different model layers. Based on this finding, we propose Surgical Steering, a novel inference-time method that disentangles these two objectives. By applying targeted activation steering to distinct layers, our approach achieves a better balance between the two competing dimensions, effectively overcoming the limitations of current alignment techniques.

多语言模型文化敏感表征对齐推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。