用平行语料和视觉化文本提升多语言视觉模型公平评估与性能
Benchmarking and Boosting Multilingual Capabilities of LVLMs via OCR-Centric Reinforcement Learning
- 基于10种语言的严格平行语料构建公平评测基准
- 视觉输入中嵌入文本可显著降低跨语言性能差距
- 无需标注数据,通过合成OCR训练实现多语言能力提升
评估大型视觉-语言模型(LVLMs)的多语言能力仍具挑战性,因多数基准依赖非平行语料,难以区分性能差异是模型局限还是数据不一致所致。为此,我们提出PM4Bench,首个基于严格平行10语言语料的多模态、多语言、多任务基准,支持模型性能的公平直接比较。我们还设计了一种将文本输入直接嵌入图像的视觉设置,更贴近LVLM代理在虚拟或物理环境中通过统一视觉输入交互的实际场景。对10个LVLM的实验表明,当文本以视觉形式呈现时,光学字符识别(OCR)是导致跨语言差异的关键因素。受此启发,我们设计了一种基于全合成、无标签OCR数据的OCR中心化GRPO训练策略,无需昂贵的任务特定VQA监督。该方法显著提升了模型的通用多语言VQA能力,在视觉设置下减少了跨语言差异,并在超出PM4Bench的场景中实现迁移增益。该方法为实现更公平的多语言视觉模型部署提供了一条高效、无标签的路径。
原文摘要 · Abstract (English)
Evaluating the multilingual capabilities of Large Vision-Language Models (LVLMs) remains challenging because most benchmarks rely on non-parallel corpora, making it unclear whether cross-lingual performance gaps reflect model limitations or dataset inconsistencies. To address this, we introduce PM4Bench, the first multimodal, multilingual, multi-task benchmark built on a strictly parallel 10-language corpus, enabling fair, apples-to-apples cross-lingual comparison of model performance. We further introduce a vision setting that embeds textual inputs directly into images, better approximating deployment scenarios where LVLM-driven agents interact with virtual or physical environments through unified visual observations. Experiments with 10 LVLMs reveal that OCR is a key factor behind cross-lingual disparity when textual content is rendered visually. Motivated by this, we design an OCR-centric GRPO training strategy using fully synthesized, label-free OCR data, without expensive task-specific VQA supervision. The resulting model improves general multilingual VQA capability, reduces cross-lingual disparities under the vision setting, and transfers gains beyond PM4Bench. This methodology offers an efficient, label-free pathway toward more equitable multilingual deployment of LVLM-driven agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。