arXiv:2510.07213cs.CLcs.AI2025-10Conference of the …被引 2

发现大模型跨语言转换依赖稀疏维度,可低成本精准控制输出语言。

Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models

  • 定位中间层到输出层中稳定存在的稀疏语言维度,无需训练。
  • 仅需50句数据即可实现跨语言切换,性能优于传统神经元方法。
  • 适合需要高效多语言控制的场景,如翻译、内容生成系统。

大型语言模型虽仅接触少量非英语数据,却具备强大多语言能力。已有研究发现,以英语为中心的大模型在中间层将多语言内容映射为英语对齐表示,并在最后一层投影回目标语言词空间。基于此,我们假设这一跨语言转换由一组固定索引的稀疏维度主导。据此提出一种无需训练的简单方法,仅需50句平行或单语数据即可识别并操控这些维度。多语言生成控制实验表明,干预这些维度可有效切换输出语言且保持语义一致,性能显著优于先前基于神经元的方法,成本却大幅降低。

原文摘要 · Abstract (English)

Large language models exhibit strong multilingual capabilities despite limited exposure to non-English data. Prior studies show that English-centric large language models map multilingual content into English-aligned representations at intermediate layers and then project them back into target-language token spaces in the final layer. From this observation, we hypothesize that this cross-lingual transition is governed by a small and sparse set of dimensions, which occur at consistent indices across the intermediate to final layers. Building on this insight, we introduce a simple, training-free method to identify and manipulate these dimensions, requiring only as few as 50 sentences of either parallel or monolingual data. Experiments on a multilingual generation control task reveal the interpretability of these dimensions, demonstrating that the interventions in these dimensions can switch the output language while preserving semantic content, and that it surpasses the performance of prior neuron-based approaches at a substantially lower cost.

多语言控制稀疏维度高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。