arXiv:2509.19563cs.CLcs.LG2025-09

分析18种语言下像素语言模型的不确定性,发现拉丁语不确定性更低

Uncertainty in Semantic Language Modeling with PIXELS

  • 用蒙特卡洛丢弃等方法量化像素模型在多语言中的不确定性
  • 像素模型在补丁重建时低估了不确定性,拉丁语系表现更稳定
  • 集成学习结合超参数调优,在16种语言任务中提升命名实体识别效果

基于像素的语言模型旨在解决传统语言建模中的词汇瓶颈问题,但不确定性量化仍是未解难题。本文针对18种语言、7种书写系统,在3个语义挑战性任务中分析了像素语言模型的不确定性与置信度。通过蒙特卡洛丢弃、Transformer注意力机制和集成学习等方法实现分析。结果表明,像素模型在补丁重建时普遍低估不确定性,且不确定性受书写系统影响:拉丁字母语言表现出更低的不确定性。在命名实体识别和问答任务中,结合超参数调优的集成学习方法在16种语言上取得更优性能。

原文摘要 · Abstract (English)

Pixel-based language models aim to solve the vocabulary bottleneck problem in language modeling, but the challenge of uncertainty quantification remains open. The novelty of this work consists of analysing uncertainty and confidence in pixel-based language models across 18 languages and 7 scripts, all part of 3 semantically challenging tasks. This is achieved through several methods such as Monte Carlo Dropout, Transformer Attention, and Ensemble Learning. The results suggest that pixel-based models underestimate uncertainty when reconstructing patches. The uncertainty is also influenced by the script, with Latin languages displaying lower uncertainty. The findings on ensemble learning show better performance when applying hyperparameter tuning during the named entity recognition and question-answering tasks across 16 languages.

语言建模不确定性像素模型多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。