arXiv:2606.17234cs.CL2026-06

让大模型自己说出翻译信心,比看内部数据更可靠。

Speaking in Self-Assessing Tongues: On the Verbalized Confidence of LLMs in Machine Translation

论文配图:Speaking in Self-Assessing Tongues: On the Verbalized Confidence of LLMs in Machine Translation
图 1 · 摘自论文原文
  • 设计五种让模型自评每个词信心的说话方法。
  • 自评信心与真实错误匹配度高,不依赖内部数据。
  • 适合评估翻译模型可靠性,尤其关注细粒度错误。

大型语言模型在机器翻译中的广泛应用,亟需对其自身输出可信度进行深入研究。与多数生成任务不同,翻译错误和信心水平可在词、字或片段等不同粒度上衡量。基于内部信号(如预测概率)的无监督方法可能误导,因其反映的是选项间的确定性而非正确性,且需访问内部数据。本文提出五种无需内部信号的模型自述信心方法,并与模型内部确定性信号的可靠性进行对比。通过细粒度错误检测和校准两个维度评估可靠性,结果表明,内部与自述方法表现相当,但不同模型间存在差异。有趣的是,内部信心与自述信心之间相关性极低甚至无相关性。

原文摘要 · Abstract (English)

The rapid rise in popularity of large language models (LLMs) for translation calls for a thorough study of the reliability of their confidence in their own outputs. Unlike many generation tasks, translation errors and confidence levels can be useful at different levels of granularity (tokens, words, or spans). Unsupervised approaches based on internal signals like predicted probabilities can be misleading because they reflect certainty among alternatives rather than correctness. In addition, they require access to such internal signals. Here, we devise five verbalized methods of extracting an LLM's per-token confidence without those shortcomings and compare their reliability with that of the model's internal signals of certainty. We evaluate reliability using two forms of alignment: fine-grained error detection and calibration. For both, internal and verbalized methods perform similarly, although results vary by model. Interestingly, we find little to no correlation between internal and verbalized methods.

大模型翻译评估置信度自省

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。