构建高质量多语言图文翻译数据集与评估体系,提升模型跨语言理解能力。
Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation
- 提出AibTrans数据集,人工校验并修正OCR错误,保障语义文化准确性。
- 对比17个模型发现:端到端模型更依赖OCR识别,生成行为差异显著。
- 设计DA Score评估指标,解决复杂场景下评价可靠性问题,适合研究者使用。
视觉语言翻译(VLT)是一项挑战性任务,要求准确识别图像中嵌入的多语言文本,并在视觉上下文支持下将其翻译为目标语言。尽管近期大视觉语言模型(LVLMs)展现出强大的多语言和视觉理解能力,但其在VLT任务上的系统性评估与理解仍显不足。本文从数据质量、模型架构和评估指标三个关键角度对VLT展开全面研究:(1) 识别现有数据集在语义与文化保真度上的关键缺陷,提出AibTrans——一个经过人工验证、带OCR修正标注的多语言平行数据集;(2) 在端到端与级联架构上对11个商用LVLM/LLM及6个开源先进模型进行基准测试,揭示其对OCR的依赖性,并对比生成与推理行为差异;(3) 提出密度感知评估方法,解决不同上下文复杂度下的评估可靠性问题,引入DA Score作为更稳健的翻译质量衡量标准。基于上述发现,我们建立新的VLT评估基准。值得注意的是,我们在高资源语言对上微调会损害跨语言性能,因此提出一种平衡的多语言微调策略,有效适配LVLM至VLT任务而不牺牲泛化能力。
原文摘要 · Abstract (English)
Vision-Language Translation (VLT) is a challenging task that requires accurately recognizing multilingual text embedded in images and translating it into the target language with the support of visual context. While recent Large Vision-Language Models (LVLMs) have demonstrated strong multilingual and visual understanding capabilities, there is a lack of systematic evaluation and understanding of their performance on VLT. In this work, we present a comprehensive study of VLT from three key perspectives: data quality, model architecture, and evaluation metrics. (1) We identify critical limitations in existing datasets, particularly in semantic and cultural fidelity, and introduce AibTrans -- a multilingual, parallel, human-verified dataset with OCR-corrected annotations. (2) We benchmark 11 commercial LVLMs/LLMs and 6 state-of-the-art open-source models across end-to-end and cascaded architectures, revealing their OCR dependency and contrasting generation versus reasoning behaviors. (3) We propose Density-Aware Evaluation to address metric reliability issues under varying contextual complexity, introducing the DA Score as a more robust measure of translation quality. Building upon these findings, we establish a new evaluation benchmark for VLT. Notably, we observe that fine-tuning LVLMs on high-resource language pairs degrades cross-lingual performance, and we propose a balanced multilingual fine-tuning strategy that effectively adapts LVLMs to VLT without sacrificing their generalization ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。