arXiv:2604.07095cs.CL2026-04

多语言嵌入模型无法跨语料库通用判断语言水平。

Multilingual Embedding Probes Fail to Generalize Across Learner Corpora

  • 用线性与非线性探测器分析模型隐藏层,预测学习者文本的CEFR等级。
  • 同语料内表现良好(加权卡帕约0.7),但跨语料性能骤降。
  • 模型学的是语料特有特征,而非通用语言能力,适合研究者参考。

多语言嵌入模型是否编码了通用的语言能力表示?我们通过在七个嵌入模型(0.3-8B)的隐藏状态上训练线性与非线性探测器,预测九个语料库、七种语言中学习者文本的CEFR等级。对比五种探测架构与基于表层文本特征的基线模型。在分布内评估中,探测器表现优异(加权卡帕约0.7),显著优于基线,中间层预测效果最佳。然而在跨语料评估中,所有探测器与模型规模下性能均崩溃。残差分析显示,跨语料探测器收敛于均匀标签预测,表明所学映射捕捉的是语料特有属性(主题、语言、任务类型、评分方法),而非抽象可迁移的能力维度。结果表明当前多语言嵌入未直接编码通用语言能力,对基于表示的自适应语言技术有重要启示。

原文摘要 · Abstract (English)

Do multilingual embedding models encode a language-general representation of proficiency? We investigate this by training linear and non-linear probes on hidden-state activations from seven embedding models (0.3-8B) to predict CEFR proficiency levels from learner texts across nine corpora and seven languages. We compare five probing architectures against a baseline trained on surface-level text features. Under in-distribution evaluation, probes achieve strong performance (Quadratic Weighted Kappa $\approx0.7$), substantially outperforming the surface baseline, with middle layers consistently yielding the best predictions. However, in cross-corpus evaluation performance collapses across all probe types and model sizes. Residual analysis reveals that out-of-distribution probes converge towards predicting uniformly distributed labels, indicating that the learned mappings capture corpus-specific distributional properties (topic, language, task type, rating methodology) rather than an abstract, transferable proficiency dimension. These results suggest that current multilingual embeddings do not straightforwardly encode language-general proficiency, with implications for representation-based approaches to proficiency-adaptive language technology.

多语言嵌入评测语言能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。