arXiv:2410.03277cs.CL2024-10被引 1

针对用户生成内容翻译中的情感与俚语难题,提出多任务评估框架。

A Multi-task Learning Framework for Evaluating Machine Translation of Emotion-loaded User-generated Content

  • 构建多任务学习架构,同步评估翻译质量与情感
  • 在多个数据集上达到当前最优性能,超越主流方法
  • 适合需要精准评估社交文本翻译的NLP研究者

用户生成内容(UGC)的机器翻译面临俚语、情感及讽刺等表达方式的挑战,现有评价指标难以覆盖这些特征。为此,我们基于已有的情感标注数据集,补充句子级评分与词级标签,构建适用于句级与词级翻译评估及情感分类的多任务数据集。提出一种新架构,结合纳什损失与对齐损失等启发式策略,设计联合损失函数,实现多任务并行处理。通过消融实验对比多种微调与多任务学习方法,验证了模型在多个数据集上的泛化能力。结果表明,该方法在翻译评估任务中达到最新最佳表现,并提供对UGC机器翻译评估的全面分析。

原文摘要 · Abstract (English)

Machine translation (MT) of user-generated content (UGC) poses unique challenges, including handling slang, emotion, and literary devices like irony and sarcasm. Evaluating the quality of these translations is challenging as current metrics do not focus on these ubiquitous features of UGC. To address this issue, we utilize an existing emotion-related dataset that includes emotion labels and human-annotated translation errors based on Multi-dimensional Quality Metrics. We extend it with sentence-level evaluation scores and word-level labels, leading to a dataset suitable for sentence- and word-level translation evaluation and emotion classification, in a multi-task setting. We propose a new architecture to perform these tasks concurrently, with a novel combined loss function, which integrates different loss heuristics, like the Nash and Aligned losses. Our evaluation compares existing fine-tuning and multi-task learning approaches, assessing generalization with ablative experiments over multiple datasets. Our approach achieves state-of-the-art performance and we present a comprehensive analysis for MT evaluation of UGC.

机器翻译情感识别多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。