arXiv:2608.19003cs.CL2026-08

研究发现,现有表示统计量无法可靠判断非洲语言任务的难易,不适用于自适应推理。

Structure, Association, and Decision Value: Representation-Based Difficulty Estimation for Adaptive Inference in African-Language NLI

  • 通过分析多语言表示统计量,评估其作为难度信号的可行性
  • 大模型在部分语言表现更好,但整体无显著优势,且存在测试数据泄露
  • 不同指标对决策效果预测能力差异大,单一信号难指导高效推理

我们探讨内部表示统计量能否为多语言非洲语种自然语言推理中的自适应推理提供有效的例级难度信号,结果表明在此设定下不可行。基于冻结的现成检查点,在15种非洲语言的AfriXNLI上报告四项发现:首先,AfriXNLI的英语配置与XNLI评测数据有1047/1050个例子完全一致,一个常用NLI检查点在该测试集上得分1.000,与XNLI测试暴露一致;因此,其英语、法语和斯瓦希里语配置无法作为对XNLI训练模型的干净评估。其次,参数量无法可靠排序各语言能力:我们的大检查点在7种语言中表现更优,但在8种中更差,总体无显著差异。第三,在三个多语言表示空间中,角度分散度始终比有效秩更受语言影响,因此合并相关性可能夸大其一而掩盖另一。第四,经语言控制后仍有效的关联取决于目标:有效秩可预测升级带来的概率提升,但不能预测是否改变预测;而廉价模型置信度则相反;两者相关系数仅为0.655。在所测试模型、信号和计算预算下,无任何评估信号使自适应路由优于始终使用昂贵模型的推理,尽管理想情况(即“上帝视角”)可在60%算力下提升11个准确率点。核心方法论发现是:一个表示统计量可能在某一类计算收益上统计显著,却对另一类无关,因而不适合作为决策变量。

原文摘要 · Abstract (English)

We ask whether internal representation statistics can provide useful example-level difficulty signals for adaptive inference in multilingual African NLP, and find that they cannot in this setting. Studying natural language inference across 15 African languages with frozen off-the-shelf checkpoints, we report four results. First, AfriXNLI's English configuration shares 1,047 of its 1,050 examples verbatim with XNLI evaluation data, and one widely used NLI checkpoint scores 1.000 on that test split, consistent with XNLI test exposure. Because AfriXNLI is derived from XNLI, its English, French and Swahili configurations cannot serve as clean evaluations for XNLI-trained models. Second, parameter count does not reliably order capability across African languages: our larger checkpoint is better in seven languages and worse in eight, with no significant aggregate difference. Third, across three multilingual representation spaces, angular dispersion is consistently more language-determined than effective rank, so pooled correlations can inflate one and mask the other. Fourth, the association that survives language control depends on the target: effective rank predicts probability gain from escalation but not whether escalation changes the prediction, while cheap-model confidence shows the opposite pattern; the two targets correlate at only 0.655. Under the tested models, signals, and compute budgets, no evaluated signal makes adaptive routing preferable to always-expensive inference, although an oracle exceeds it by 11 accuracy points at 60% of the compute. Our central methodological finding is that a representation statistic can be statistically significant for one notion of computational benefit while being irrelevant to another, and therefore be a poor decision variable.

自然语言推理非洲语言自适应推理表示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。