arXiv:2412.13717cs.CLcs.CV2024-12NAACL被引 6

提出一套自动评估图像跨文化转译效果的指标体系。

Towards Automatic Evaluation for Image Transcreation

  • 基于机器翻译思路设计三类自动评估指标:对象、嵌入和视觉语言模型。
  • 在7个国家的元评估中,指标与人工评分相关性达0.55–0.87。
  • 适用于需跨文化适配视觉内容的AI系统开发与评测。

超越传统语音与文本翻译范式,近年来自动化图像跨文化转译(image transcreation)成为视觉内容跨文化传播的新方向。然而,该任务因缺乏自动评估机制而难以形式化,此前研究均依赖人工评价。本文提出一套受机器翻译评估启发的自动指标体系,分为三类:基于对象、基于嵌入、基于视觉语言模型(VLM)。结合翻译学理论与实际转译实践,识别出图像转译的三大核心维度:文化相关性、语义等价性和视觉相似性,并设计对应指标进行评估。实验表明,专有VLM在捕捉文化相关性与语义等价性上表现最佳,而视觉编码器表征擅长衡量视觉相似性。在7个国家的元评估中,各指标与人工评分的相关性达到0.55–0.87,整体一致性较高。最后,通过分析各类指标优劣,构建了一个兼具理论基础与实用价值的自动化图像转译评估框架。代码已公开:https://github.com/simran-khanuja/automatic-eval-img-transcreation。

原文摘要 · Abstract (English)

Beyond conventional paradigms of translating speech and text, recently, there has been interest in automated transcreation of images to facilitate localization of visual content across different cultures. Attempts to define this as a formal Machine Learning (ML) problem have been impeded by the lack of automatic evaluation mechanisms, with previous work relying solely on human evaluation. In this paper, we seek to close this gap by proposing a suite of automatic evaluation metrics inspired by machine translation (MT) metrics, categorized into: a) Object-based, b) Embedding-based, and c) VLM-based. Drawing on theories from translation studies and real-world transcreation practices, we identify three critical dimensions of image transcreation: cultural relevance, semantic equivalence and visual similarity, and design our metrics to evaluate systems along these axes. Our results show that proprietary VLMs best identify cultural relevance and semantic equivalence, while vision-encoder representations are adept at measuring visual similarity. Meta-evaluation across 7 countries shows our metrics agree strongly with human ratings, with average segment-level correlations ranging from 0.55-0.87. Finally, through a discussion of the merits and demerits of each metric, we offer a robust framework for automated image transcreation evaluation, grounded in both theoretical foundations and practical application. Our code can be found here: https://github.com/simran-khanuja/automatic-eval-img-transcreation.

图像转译自动评估跨文化VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。