arXiv:2609.08171cs.CL2026-09

EviSI用大模型评估同传翻译,更准判断语义和口语质量。

EviSI: An Evaluation Agent for Simultaneous Interpreting

论文配图:EviSI: An Evaluation Agent for Simultaneous Interpreting
图 1 · 摘自论文原文
  • 基于MQM原则构建共享源证据,定量评估语义忠实度与口语表达
  • 中文到英文评估中与人工排名相关性达0.467,英译中达0.707
  • 适合评估同传系统,尤其关注语义保真与自然口语输出

同步语音到语音翻译需在源流持续输入时完成理解、翻译和口头表达。为保证及时输出并控制延迟,系统采用重述与摘要策略,可在保留语义的同时偏离原文。传统指标如BLEU和COMET难以区分此类变化与语义丢失。我们提出EviSI,一个基于大语言模型的评估代理,借鉴多维质量评估(MQM)的错误分析与惩罚机制。EviSI构建共享源证据,评估语义忠实度与口语表达,整合重叠错误并实现确定性评分。在英译中任务中,EviSI成功恢复了人工系统排序;在相同语料库内,其平均肯德尔相关系数达到0.707(英→中)和0.467(中→英),优于对比基线。跨五个方向扩展验证显示,其结果与COMET具正向一致性,无需人工标注。但个体输出与人类评价仍存在混合一致性。

原文摘要 · Abstract (English)

Simultaneous speech-to-speech translation requires understanding, translation and spoken delivery while the source stream continues. To support timely delivery and limit accumulated delay, systems adopt reformulation and summarization, which can preserve meaning while departing from written references. BLEU and COMET may not reliably distinguish such variation from semantic loss. We introduce EviSI, a large language model evaluation agent adapting the error analysis and penalty principles of Multidimensional Quality Metrics (MQM). It constructs shared source evidence, assesses semantic fidelity and oral expression, reconciles overlapping errors and scores deterministically. EviSI recovers the aggregate human system ranking for English to Chinese. Mean Kendall agreement with human system rankings within corpora reaches 0.707 for English to Chinese and 0.467 for Chinese to English, exceeding evaluated baselines. An extension across five directions shows positive concordance with COMET without human ratings. Individual output agreement with humans remains mixed.

同传评估大模型评测语义保真口语质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。