arXiv:2507.02357cs.CL2025-07中稿 · ACL

通过检索相似样例与置信度加权融合,提升科学视觉问答性能。

Coling-UniA at SciVQA 2025: Few-Shot Example Retrieval and Confidence-Informed Ensembling for Multimodal Large Language Models

  • 基于问题类型和图表特征选择模型与少样本示例
  • 在盲测数据上平均F1达85.12,排名第三
  • 适合需要高精度科学图文理解的场景

本文介绍了我们在SciVQA 2025科学视觉问答共享任务中的系统。系统采用两个多模态大模型的集成,并结合多种少样本示例检索策略。模型及少样本设置根据图表类型和问题类型进行选择。同时,根据模型预测置信度筛选答案。在盲测数据上,系统在7个参赛者中排名第三,平均F1分数为85.12(基于ROUGE-1、ROUGE-L和BERTS)。代码已公开。

原文摘要 · Abstract (English)

This paper describes our system for the SciVQA 2025 Shared Task on Scientific Visual Question Answering. Our system employs an ensemble of two Multimodal Large Language Models and various few-shot example retrieval strategies. The model and few-shot setting are selected based on the figure and question type. We also select answers based on the models' confidence levels. On the blind test data, our system ranks third out of seven with an average F1 score of 85.12 across ROUGE-1, ROUGE-L, and BERTS. Our code is publicly available.

视觉问答多模态少样本学习模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。