首个多语言模型行为操控基准,可系统评估跨语言控制效果。
CLaS-Bench: A Cross-Lingual Alignment and Steering Benchmark
- 构建32语言平行问题集,量化模型语言控制能力
- 残差向量干预法在各语言中表现最优,平均得分领先15%以上
- 揭示语言特征主要存在于深层,且按语系聚类,适合低资源适配研究
理解与控制大语言模型在多语言场景下的行为日益重要。除提示或微调外,通过推理阶段操纵内部表征的“操控”技术因其高效与可解释性成为新趋势。然而,尚无专门的评估基准来衡量操控方法的有效性。本文提出CLaS-Bench,一个面向32种语言的轻量级平行问题基准,用于系统评估多语言操控方法。我们评测了包括残差流DiffMean干预、探测导出方向、语言特异性神经元、PCA/LDA向量、稀疏自编码器及提示基线在内的多种方法。操控性能从语言控制力与语义相关性两个维度评估,并以调和均值综合得分。结果表明,简单残差式DiffMean方法在所有语言中持续优于其他方法;层分析显示语言特异性结构主要出现在深层,且操控方向按语言家族聚类。CLaS-Bench是首个标准化的多语言操控评估基准,既支持语言表征的科学分析,也提供低成本适配的实用评估工具。
原文摘要 · Abstract (English)
Understanding and controlling the behavior of large language models (LLMs) is an increasingly important topic in multilingual NLP. Beyond prompting or fine-tuning, , i.e.,~manipulating internal representations during inference, has emerged as a more efficient and interpretable technique for adapting models to a target language. Yet, no dedicated benchmarks or evaluation protocols exist to quantify the effectiveness of steering techniques. We introduce CLaS-Bench, a lightweight parallel-question benchmark for evaluating language-forcing behavior in LLMs across 32 languages, enabling systematic evaluation of multilingual steering methods. We evaluate a broad array of steering techniques, including residual-stream DiffMean interventions, probe-derived directions, language-specific neurons, PCA/LDA vectors, Sparse Autoencoders, and prompting baselines. Steering performance is measured along two axes: language control and semantic relevance, combined into a single harmonic-mean steering score. We find that across languages simple residual-based DiffMean method consistently outperforms all other methods. Moreover, a layer-wise analysis reveals that language-specific structure emerges predominantly in later layers and steering directions cluster based on language family. CLaS-Bench is the first standardized benchmark for multilingual steering, enabling both rigorous scientific analysis of language representations and practical evaluation of steering as a low-cost adaptation alternative.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。