评估多语言图像描述生成效果,发现CLIP模型跨语言表现良好。
Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?
- 用机器翻译数据+人工评分构建多语言评测集
- 微调后的多语言CLIPScore与人工评分高度相关
- 适合关注跨语言视觉-语言理解的研究者
图像描述评估在语言流畅性和语义与视觉内容对应性方面已有显著进展。尽管出现了CLIPScore等新指标,多语言描述评估仍相对不足。本文提出多种策略并开展广泛实验,研究CLIPScore在多语言环境下的表现。为解决多语言测试数据缺乏问题,采用两种方法:(1) 使用带有质量感知的机器翻译数据集及人工评分;(2) 重新利用面向语义推理和理解的多语言数据集。结果表明,微调后的多语言模型具有良好的跨语言泛化能力,可应对复杂语言挑战。在机器翻译数据上的测试显示,多语言CLIPScore模型在不同语言间保持与人工评分的高相关性;在原生多语言、多文化数据上的额外测试进一步验证了其高质量评估能力。
原文摘要 · Abstract (English)
The evaluation of image captions, looking at both linguistic fluency and semantic correspondence to visual contents, has witnessed a significant effort. Still, despite advancements such as the CLIPScore metric, multilingual captioning evaluation has remained relatively unexplored. This work presents several strategies, and extensive experiments, related to evaluating CLIPScore variants in multilingual settings. To address the lack of multilingual test data, we consider two different strategies: (1) using quality aware machine-translated datasets with human judgements, and (2) re-purposing multilingual datasets that target semantic inference and reasoning. Our results highlight the potential of finetuned multilingual models to generalize across languages and to handle complex linguistic challenges. Tests with machine-translated data show that multilingual CLIPScore models can maintain a high correlation with human judgements across different languages, and additional tests with natively multilingual and multicultural data further attest to the high-quality assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。