arXiv:2507.11114cs.CL2025-07被引 1

轻量级多模态模型集成在多语言推理中表现卓越,夺冠并领跑11个语种赛道。

MSA at ImageCLEF 2025 Multimodal Reasoning: Multilingual Multimodal Reasoning With Ensemble Vision Language Models

  • 用Gemini系列模型分步处理视觉描述、摘要优化和最终推理,通过精心设计提示协同工作。
  • 在官方榜单上以81.4%准确率获多语言赛道第一,13个语种中11个领先,克罗地亚达95.07%。
  • 零样本下Gemini 2.5 Flash优于训练模型,提示工程显著提升准确率(55.9%→61.7%)。

我们提出一个稳健的集成式系统,用于ImageCLEF 2025 EXAMS V挑战中的多语言多模态推理任务。该系统结合Gemini 2.5 Flash生成视觉描述,Gemini 1.5 Pro进行摘要优化与一致性检查,以及Gemini 2.5 Pro作为推理器完成最终答案选择,所有模块通过精心设计的少样本与零样本提示协调。我们进行了广泛消融实验,在英文数据集及其多语言增强版本上训练了多个大语言模型(Gemini 2.5 Flash、Phi 4、Gemma 3、Mistral)。此外,评估了Gemini 2.5 Flash在零样本设置下的表现,发现其显著优于训练模型。提示设计亦至关重要:强制使用简洁、语言标准化格式并禁止解释性文本,使英文验证集准确率从55.9%提升至61.7%。在官方排行榜上,我们的系统(团队MSA)在多语言赛道以81.4%准确率获得第一名,并在13个语种中的11个领先,包括克罗地亚语95.07%和意大利语92.12%的优异成绩。结果表明,轻量级OCR-VLM集成搭配精准提示策略与跨语言数据增强,可在高要求的多语言教育场景中超越更复杂的端到端模型。

原文摘要 · Abstract (English)

We present a robust ensemble-based system for multilingual multimodal reasoning, designed for the ImageCLEF 2025 EXAMS V challenge. Our approach integrates Gemini 2.5 Flash for visual description, Gemini 1.5 Pro for caption refinement and consistency checks, and Gemini 2.5 Pro as a reasoner which handles final answer selection, all coordinated through carefully engineered few-shot and zero-shot prompts. We conducted an extensive ablation study, training several large language models (Gemini 2.5 Flash, Phi 4, Gemma 3, Mistral) on an English dataset and its multilingual augmented version. Additionally, we evaluated Gemini 2.5 Flash in a zero-shot setting for comparison and found it to substantially outperform the trained models. Prompt design also proved critical: enforcing concise, language-normalized formats and prohibiting explanatory text boosted model accuracy on the English validation set from 55.9% to 61.7%. On the official leaderboard, our system (Team MSA) achieved first place overall in the multilingual track with 81.4% accuracy, and led 11 out of 13 individual language tracks, with top results such as 95.07% for Croatian and 92.12% for Italian. These findings highlight that lightweight OCR-VLM ensembles, when paired with precise prompt strategies and cross-lingual augmentation, can outperform heavier end-to-end models in high-stakes, multilingual educational settings.

多模态推理多语言模型集成提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。