arXiv:2411.02589cs.CL2024-11被引 4

用多模态大模型提升漫画翻译质量,融合图像信息解决语义歧义。

Context-Informed Machine Translation of Manga using Multimodal Large Language Models

  • 利用多模态大模型的视觉能力,结合图像上下文优化翻译。
  • 在日英翻译上达到顶尖水平,日波翻译创历史新高。
  • 开源工具与首个日波漫画平行语料库,助力后续研究。

由于手工翻译耗时耗力,大多数漫画仍局限于日本国内市场。自动漫画翻译是潜在解决方案,但该领域仍处于萌芽阶段,且复杂度高于常规翻译,需有效整合视觉元素以消除歧义。本文探究多模态大语言模型(MLLMs)在漫画翻译中的应用潜力,提出一种利用视觉组件提升翻译质量的方法,评估了翻译单元大小、上下文长度的影响,并提出一种高效的令牌处理策略。此外,我们构建了首个平行的日语-波兰语漫画翻译数据集,作为未来研究的基准。最后,我们开源了一套软件工具,支持他人对MLLMs在漫画翻译中的表现进行评测。实验结果表明,所提方法在日英翻译中达到当前最优性能,日波翻译亦创下新纪录。

原文摘要 · Abstract (English)

Due to the significant time and effort required for handcrafting translations, most manga never leave the domestic Japanese market. Automatic manga translation is a promising potential solution. However, it is a budding and underdeveloped field and presents complexities even greater than those found in standard translation due to the need to effectively incorporate visual elements into the translation process to resolve ambiguities. In this work, we investigate to what extent multimodal large language models (LLMs) can provide effective manga translation, thereby assisting manga authors and publishers in reaching wider audiences. Specifically, we propose a methodology that leverages the vision component of multimodal LLMs to improve translation quality and evaluate the impact of translation unit size, context length, and propose a token efficient approach for manga translation. Moreover, we introduce a new evaluation dataset -- the first parallel Japanese-Polish manga translation dataset -- as part of a benchmark to be used in future research. Finally, we contribute an open-source software suite, enabling others to benchmark LLMs for manga translation. Our findings demonstrate that our proposed methods achieve state-of-the-art results for Japanese-English translation and set a new standard for Japanese-Polish.

漫画翻译多模态大模型NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。