arXiv:2607.03936cs.CL2026-07

不微调模型,也能通过激活特定神经元或向量方向控制阿拉伯语模型生成方言。

Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs

论文配图:Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs
图 1 · 摘自论文原文
  • 定位稀疏神经元,通过激活/抑制实现方言引导。
  • 提取方言特异性激活方向,推理时注入以提升方言准确性。
  • 无需额外训练,适合资源少的方言生成任务。

阿拉伯语自然语言处理面临方言数据远少于现代标准阿拉伯语(MSA)的挑战,导致大模型过度生成MSA且难以准确生成方言。本文从可解释性视角出发,探究方言特征在模型内部的编码位置与方式,并提出两种无需微调的推理阶段控制方法。首先,通过神经元层面分析,识别出编码方言特性的稀疏神经元群,证明增强或抑制这些神经元可有效引导模型输出至目标方言。其次,针对单个神经元中方言特征高度纠缠的问题,采用向量引导方法,提取方言特异的激活方向并在推理时注入。两种方法共同揭示了阿拉伯语大模型中方言知识的几何结构,构建了一个基于可解释性的、无需方言微调的方言控制框架。

原文摘要 · Abstract (English)

A key challenge in Arabic NLP is the scarcity of dialectal data relative to Modern Standard Arabic (MSA), causing LLMs to overproduce MSA and struggle with dialectally accurate generation. From an interpretability perspective, this raises a fundamental question: where and how are dialectal features encoded within model internals, and can these representations be leveraged to improve dialect generation without fine-tuning? This study investigates two complementary inference-time approaches that serve simultaneously as interpretability probes and control mechanisms. First, we conduct a neuron-level analysis, identifying sparse neuron populations that encode dialect-specific features and showing that amplifying or suppressing these neurons can steer model outputs toward target dialects. Second, motivated by the entanglement of dialectal features at the single-neuron level, we apply a vector-steering approach that extracts dialect-specific activation directions and injects them during inference. Together, these methods illuminate the geometry of dialectal knowledge in Arabic LLMs and offer a principled, interpretability-grounded framework for dialect control without requiring dialect-specific fine-tuning.

阿拉伯语方言生成神经元控制向量注入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。