arXiv:2410.10861cs.CL2024-10EMNLP被引 1

可视化工具帮助定位翻译模型错误并分析缺陷原因

Translation Canvas: An Explainable Interface to Pinpoint and Analyze Translation Systems

  • 通过多维度指标识别系统级错误模式与严重程度
  • 支持逐句错误定位,显示错误片段及解释说明
  • 相比COMET和SacreBLEU更易用、更易理解

随着机器翻译研究的快速发展,评估工具已成为衡量系统进展的关键。COMET和SacreBLEU等工具虽能提供单个质量分数,适用于系统间对比,但对细粒度系统性能比较和实例级错误分析支持有限。为此,我们提出Translation Canvas——一种可解释的交互界面,用于精准定位和分析翻译系统表现:1)帮助研究人员理解系统级模型性能,识别常见错误(频率与严重性),并基于多种评估指标分析不同系统间的关系;2)支持细粒度分析,通过高亮错误片段并提供解释,同时可选择性展示各系统的预测结果。人工评估表明,Translation Canvas在可读性和易用性上优于COMET和SacreBLEU。

原文摘要 · Abstract (English)

With the rapid advancement of machine translation research, evaluation toolkits have become essential for benchmarking system progress. Tools like COMET and SacreBLEU offer single quality score assessments that are effective for pairwise system comparisons. However, these tools provide limited insights for fine-grained system-level comparisons and the analysis of instance-level defects. To address these limitations, we introduce Translation Canvas, an explainable interface designed to pinpoint and analyze translation systems' performance: 1) Translation Canvas assists machine translation researchers in comprehending system-level model performance by identifying common errors (their frequency and severity) and analyzing relationships between different systems based on various evaluation metrics. 2) It supports fine-grained analysis by highlighting error spans with explanations and selectively displaying systems' predictions. According to human evaluation, Translation Canvas demonstrates superior performance over COMET and SacreBLEU packages under enjoyability and understandability criteria.

机器翻译可解释性评估工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。