评测聊天翻译系统在多语言对话中的表现,发现单句翻译好但整体对话质量仍有提升空间。
Findings of the WMT 2024 Shared Task on Chat Translation
- 基于对话上下文的多语言客服对话翻译任务,含英韩、英荷等新语对。
- 22个主提交+32个对比提交,每语对至少3支团队参与,以人工评分定排名。
- 系统擅长单轮翻译,但跨轮对话连贯性与整体质量仍待改进。
本文报告了第三届聊天翻译共享任务的成果。该任务延续以往形式,聚焦于双语客户支持对话的翻译,特别关注对话上下文对翻译质量和评估的影响。本次新增英语-韩语、英语-荷兰语两个语对,保留此前的英语-德语、英语-法语及英语-巴西葡萄牙语。共收到8支队伍提交的22个主系统和32个对比系统,每个语对均有至少三支团队参与。采用自动指标与人工直接评估相结合的方式进行综合评测。各语对的官方排名依据人工评估得分确定,涵盖代理与客户双向翻译表现。分析显示,尽管系统在单轮对话翻译上表现优异,但在整体对话层面的翻译质量仍有提升空间。
原文摘要 · Abstract (English)
This paper presents the findings from the third edition of the Chat Translation Shared Task. As with previous editions, the task involved translating bilingual customer support conversations, specifically focusing on the impact of conversation context in translation quality and evaluation. We also include two new language pairs: English-Korean and English-Dutch, in addition to the set of language pairs from previous editions: English-German, English-French, and English-Brazilian Portuguese. We received 22 primary submissions and 32 contrastive submissions from eight teams, with each language pair having participation from at least three teams. We evaluated the systems comprehensively using both automatic metrics and human judgments via a direct assessment framework. The official rankings for each language pair were determined based on human evaluation scores, considering performance in both translation directions--agent and customer. Our analysis shows that while the systems excelled at translating individual turns, there is room for improvement in overall conversation-level translation quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。