arXiv:2603.07330cs.CL2026-03被引 1

在噪声环境下,用不确定性估计提升多语言文本分类的可靠性。

To Predict or Not to Predict? Towards reliable uncertainty estimation in the presence of noise

  • 通过蒙特卡洛丢弃法实现稳定且鲁棒的不确定性估计。
  • 放弃最不确定的10%样本后,宏平均F1提升至0.85。
  • 适合构建真实场景下可靠的多语言NLP系统开发者。

本研究考察了不确定性估计(UE)方法在噪声和非主题条件下的多语言文本分类中的作用。基于跨多种语言的复杂与简单句子分类任务,评估了多种UE技术对不同指标的贡献,以增强预测的鲁棒性。结果表明,在高资源、领域内场景中,依赖Softmax输出的方法仍具竞争力,但在低资源或领域偏移情况下其可靠性下降。相比之下,蒙特卡洛丢弃方法在所有语言中均表现出一致的优异性能,提供了更优的校准效果、稳定的决策阈值以及更强的判别能力,即使在恶劣条件下也表现稳健。进一步证明,结合不确定性估计与可信度度量可显著提升非主题分类效果:在Readme任务中,放弃最不确定的10%实例后,宏平均F1从0.81提升至0.85。该研究为构建真实多语言环境中更可靠的NLP系统提供了可操作的洞见。

原文摘要 · Abstract (English)

This study examines the role of uncertainty estimation (UE) methods in multilingual text classification under noisy and non-topical conditions. Using a complex-vs-simple sentence classification task across several languages, we evaluate a range of UE techniques against a range of metrics to assess their contribution to making more robust predictions. Results indicate that while methods relying on softmax outputs remain competitive in high-resource in-domain settings, their reliability declines in low-resource or domain-shift scenarios. In contrast, Monte Carlo dropout approaches demonstrate consistently strong performance across all languages, offering more robust calibration, stable decision thresholds, and greater discriminative power even under adverse conditions. We further demonstrate the positive impact of UE on non-topical classification: abstaining from predicting the 10\% most uncertain instances increases the macro F1 score from 0.81 to 0.85 in the Readme task. By integrating UE with trustworthiness metrics, this study provides actionable insights for developing more reliable NLP systems in real-world multilingual environments. See https://github.com/Nouran-Khallaf/To-Predict-or-Not-to-Predict

不确定性估计多语言文本分类鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。