用多模态推理和集成模型提升科学图表问答的准确性。
Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling
- 结合链式思维与提示优化,增强对图表中数值和多步推理的理解。
- 最强模型InternVL3在测试集上达到0.740的ROUGE-1 F1和0.983的BERTScore。
- 集成多个视觉语言模型可减少错误,适合高精度科学问答场景。
技术报告和文章中常包含图表等半结构化数据,从中提取信息对问答等下游任务至关重要。现有视觉问答方法在科学数据解读上仍面临精度不足的问题,尤其体现在数值处理、多步视觉推理及图文一致性保持方面。本文针对SciVQA 2025共享任务,提出基于5B至8B参数模型的方法,用于回答源自学术文章科学图表的视觉与非视觉问题。最强单模型InternVL3在SciVQA测试集上取得ROUGE-1 F1为0.740、ROUGE-L F1为0.740、BERTScore为0.983的性能。我们还构建了融合多个视觉语言模型的集成模型,通过验证集误差分析显示其优于多数单模型,尽管InternVL3仍是最佳独立表现者。结果表明,提示优化、链式思维推理与集成建模能有效提升视觉问答能力。
原文摘要 · Abstract (English)
Technical reports and articles often contain valuable information in the form of semi-structured data like charts, and figures. Interpreting these and using the information from them is essential for downstream tasks such as question answering (QA). Current approaches to visual question answering often struggle with the precision required for scientific data interpretation, particularly in handling numerical values, multi-step reasoning over visual elements, and maintaining consistency between visual observation and textual reasoning. We present our approach to the SciVQA 2025 shared task, focusing on answering visual and non-visual questions grounded in scientific figures from scholarly articles. We conducted a series of experiments using models with 5B to 8B parameters. Our strongest individual model, InternVL3, achieved ROUGE-1 and ROUGE-L F1 scores of \textbf{0.740} and a BERTScore of \textbf{0.983} on the SciVQA test split. We also developed an ensemble model with multiple vision language models (VLMs). Through error analysis on the validation split, our ensemble approach improved performance compared to most individual models, though InternVL3 remained the strongest standalone performer. Our findings underscore the effectiveness of prompt optimization, chain-of-thought reasoning and ensemble modeling in improving the model's ability in visual question answering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。