提出评估隐喻翻译错误的框架,揭示主流模型在隐喻处理上的短板。
MetaHOPE: A Metaphor-Oriented Evaluation Framework for Analysing MT and LLM Translation Errors

- 构建面向隐喻翻译的误差严重度标注框架MetaHOPE。
- 在GoogleMT、GPT5.4、Hunyuan-7b上测试,发现隐喻翻译错误率超30%。
- 开源数据集与标注工具,适合隐喻研究和NLP模型优化者使用。
本文提出一种面向隐喻翻译的误差严重度感知标注框架MetaHOPE,用于评估机器翻译(MT)与大语言模型(LLM)在隐喻处理上的表现。隐喻具有语义复杂性、上下文依赖性和文化嵌入性,易引发自然语言理解与处理(NLU/NLP)中的歧义问题。我们选取GoogleMT、GPT5.4和Hunyuan-7b三类代表性系统,基于VUAMC与PSUCMC两个人工标注的隐喻语料库(英译中与中译英),对单语语料进行错误标注,并生成双语后编辑黄金参考文本,形成新资源。该框架及配套数据可为隐喻翻译研究提供评估基准与分析依据。相关资源已公开于github.com/Jiahui84/MetaHOPE。
原文摘要 · Abstract (English)
In this opinion paper, we propose MetaHOPE, an error severity-aware annotation framework for evaluating metaphor translations. Metaphors present challenges for machine translation (MT) and natural language understanding and processing (NLU, NLP), because it presents the features of semantic complexity, contextual dependency, and cultural embeddings that can lead to ambiguity issues for NLP models. To investigate how state-of-the-art NLP models perform on translating metaphors, we select three representative systems, i.e., GoogleMT, GPT5.4, and Hunyuan-7b as Neural MT (NMT) models and LLMs. We used two human-annotated metaphor corpora, including VUAMC and PSUCMC for English-to-Chinese and Chinese-to-English translation purposes. The original corpora we used are monolingual, where we carried out error annotation using the MetaHOPE framework, and also produced the human post-edited gold reference for bilingual use as a new resource. We believe the MetaHOPE evaluation framework for metaphor translation annotation, the parallel corpora resources, and the error analysis on SOTA automatic translation models can be useful and shed some light for the field of metaphor translation study. We share our resources publicly at github.com/Jiahui84/MetaHOPE
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。