arXiv:2506.12450cs.CL2025-06被引 4

通过隐空间注入实现多语言模型的精准跨语言控制

Language Surgery in Multilingual Large Language Models

  • 在中间层发现自然对齐表示,用于分离语言特异与通用信息
  • 提出ITLC方法,在不破坏语义的前提下实现跨语言精准控制
  • 有效缓解大模型中的语言混淆问题,适合多语言应用开发者

大型语言模型(LLMs)在任务和语言间展现出卓越的泛化能力,推动了自然语言处理的发展。本文研究了LLMs中自然出现的表征对齐现象,特别是在中间层的表现,及其对解耦语言特异性与语言无关信息的影响。我们实证验证了这种对齐的存在,分析其与显式设计对齐模型的行为差异,并展示了其在不造成语义退化的情况下实现语言特异性操控的潜力。基于这些发现,我们提出推理时语言控制(ITLC)方法,利用隐空间注入实现精确的跨语言控制,缓解大模型中的语言混淆问题。实验表明,ITLC在保持目标语言语义完整性的同时具备强大的跨语言控制能力,显著减轻了当前大规模LLMs中持续存在的语言生成不一致问题。本工作深化了对LLM表征对齐的理解,并为提升其单语与跨语言性能提供了实用解决方案。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable generalization capabilities across tasks and languages, revolutionizing natural language processing. This paper investigates the naturally emerging representation alignment in LLMs, particularly in the middle layers, and its implications for disentangling language-specific and language-agnostic information. We empirically confirm the existence of this alignment, analyze its behavior in comparison to explicitly designed alignment models, and demonstrate its potential for language-specific manipulation without semantic degradation. Building on these findings, we propose Inference-Time Language Control (ITLC), a novel method that leverages latent injection to enable precise cross-lingual language control and mitigate language confusion in LLMs. Our experiments highlight ITLC's strong cross-lingual control capabilities while preserving semantic integrity in target languages. Furthermore, we demonstrate its effectiveness in alleviating the cross-lingual language confusion problem, which persists even in current large-scale LLMs, leading to inconsistent language generation. This work advances our understanding of representation alignment in LLMs and introduces a practical solution for enhancing their monolingual and cross-lingual performance.

多语言模型语言控制表征对齐隐空间注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。