arXiv:2607.06327cs.CLcs.AI2026-07

跨语言问答中,用英文推理能显著提升模型不确定性估计能力

Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs

论文配图:Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
图 1 · 摘自论文原文
  • 让模型用英文推理,即使问题为低资源语言,也能提升不确定性判断
  • 小模型用概率方法,大模型用自述式不确定性的效果更好
  • 该研究为多语言系统如何选择性拒绝回答提供实证指导

现有不确定性估计(UE)研究主要聚焦英语,本文首次在22种语言上开展大规模评估,覆盖高、中、低资源语言。基于两个人工标注的问答数据集,比较了九种开箱与闭箱UE方法在不同模型规模和架构下的表现,通过长文本推理方式避免使用LLM作为评判者或嵌入打分带来的噪声。主要发现:第一,用英文推理而问题保持低资源语言,可显著提升UE性能,说明理解能力未受损,瓶颈在于生成;第二,英文推理能缩小高低资源语言间的UE差距,表明生成语言比问题语言更重要;第三,小模型适合概率类开箱方法,大模型则自述式不确定性更优。最后分析了选择性预测中的阈值选择,为多语言场景下的拒绝策略提供校准建议。

原文摘要 · Abstract (English)

Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 languages, spanning high-, mid-, and low-resource settings. Using two human-curated Q\&A datasets, we compare open and closed box UE methods (nine in total) across different model sizes and architectures while eliciting long-form reasoning, avoiding LLM-as-a-judge and embedding-based scoring, which can introduce evaluation noise. We report three main actionable findings. First, we find that prompting models to reason in English while keeping questions in low-resource languages substantially improves UE performance, suggesting that comprehension of low-resource languages is largely intact, and that the reliability bottleneck lies in generation rather than understanding. Second, prompting models to reason in English closes the UE performance gap between low and high-resource languages, demonstrating that generation language matters more than the question language. Third, the choice of UE method should depend on model scale: at smaller scales, open-box probability-based methods outperform alternatives; at larger scales, closed-box self-verbalized uncertainty becomes superior. Finally, we provide an analysis of threshold selection for selective prediction, offering guidance on calibrating abstention in multilingual settings.

不确定性估计多语言大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。