arXiv:2501.17187cs.CLcs.AI2025-01

用可视化工具展示大模型翻译时的置信度,提升可信度与可解释性。

Visualizing Uncertainty in Translation Tasks: An Evaluation of LLM Performance and Confidence Metrics

  • 设计三种新置信度量化指标,从概率分布角度评估翻译不确定性。
  • 发现传统评分与新指标呈线性关系,验证方法有效性。
  • 开发交互式网页工具,用颜色直观展示每个词的可信度高低。

大型语言模型(LLMs)在机器翻译中的应用日益广泛,但其预测常伴随不确定性,影响可解释性与用户信任。本文旨在实现两个目标:(1) 提供模型在词级别上的置信度洞察;(2) 构建基于Web的可视化工具以量化并呈现翻译不确定性。研究采用T5模型与WMT19数据集,使用BLEU、METEOR和ROUGE等标准指标评估翻译质量。提出三种新型不确定性量化(UQ)指标:(1) 词概率的几何平均值,(2) 词概率的算术平均值,(3) 词概率分布的峰度算术平均值。分析表明,传统评价指标与新提出的UQ指标之间存在线性关系,证明方法的有效性。此外,开发了一个交互式网页可视化工具,通过颜色渐变表示每个词的置信度水平。该工具使用户能直观理解翻译质量,并深入洞察模型表现。结果表明,所提UQ指标与可视化工具兼具鲁棒性与可解释性,为机器翻译系统的评估与使用提供了实用支持。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly utilized for machine translation, yet their predictions often exhibit uncertainties that hinder interpretability and user trust. Effectively visualizing these uncertainties can enhance the usability of LLM outputs, particularly in contexts where translation accuracy is critical. This paper addresses two primary objectives: (1) providing users with token-level insights into model confidence and (2) developing a web-based visualization tool to quantify and represent translation uncertainties. To achieve these goals, we utilized the T5 model with the WMT19 dataset for translation tasks and evaluated translation quality using established metrics such as BLEU, METEOR, and ROUGE. We introduced three novel uncertainty quantification (UQ) metrics: (1) the geometric mean of token probabilities, (2) the arithmetic mean of token probabilities, and (3) the arithmetic mean of the kurtosis of token distributions. These metrics provide a simple yet effective framework for evaluating translation performance. Our analysis revealed a linear relationship between the traditional evaluation metrics and our UQ metrics, demonstrating the validity of our approach. Additionally, we developed an interactive web-based visualization that uses a color gradient to represent token confidence. This tool offers users a clear and intuitive understanding of translation quality while providing valuable insights into model performance. Overall, we show that our UQ metrics and visualization are both robust and interpretable, offering practical tools for evaluating and accessing machine translation systems.

机器翻译不确定性可视化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。