arXiv:2509.16628cs.CV2025-09被引 3

用图文描述提升小模型科学图像问答能力

Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning

  • 结合图像描述和问答对进行监督微调
  • 在ScienceQA上显著提升小模型性能
  • 专为低资源语言设计,适合多语种研究者

本研究提出视觉-图文感知监督微调(VCASFT),一种新型学习范式,旨在提升小型视觉语言模型(VLMs)在科学视觉问答(VQA)任务上的表现。VCASFT利用图像描述作为零样本提示,结合问题-答案对与指令微调模型,实现显著性能提升。为全面评估该方法,我们在包含多语言、多学科和多领域的ScienceQA数据集上进行基准测试,验证其在多种教育场景中的适应性与有效性。此外,为验证该技术在低资源语言中的表现,我们构建了HiSciVQA,一个包含2,245个高质量、人工标注的印地语多模态问答对的数据集,填补了低资源语言问答数据集的空白,并作为测试VCASFT的基础。同时,我们引入基于大语言模型的新型评估方案,对HiSciVQA上的VLM进行评估,提供了超越传统n-gram匹配准确率的深层洞察。我们承诺开源全部代码及HiSciVQA数据集,推动学术界发展。

原文摘要 · Abstract (English)

In this study, we introduce Vision-Caption aware Supervised FineTuning (VCASFT), a novel learning paradigm designed to enhance the performance of smaller Vision Language Models(VLMs) on scientific visual question answering(VQA) tasks. VCASFT leverages image captions as zero-shot prompts alongside question-answer pairs and instruction-tunes models to yield significant performance improvements. To comprehensively evaluate VCASFT, we benchmark it on ScienceQA, which consists of questions across diverse languages, subjects, and fields, demonstrating its adaptability and effectiveness in a variety of educational contexts. Additionally, to further demonstrate the effectiveness of this technique on lowresource languages, we developed HiSciVQA, a dataset comprising 2,245 high-quality, hand-annotated Hindi multimodal Q&A pairs. This dataset addresses the critical need for low-resource language Q&A datasets and serves as a foundation for testing VCASFT. Additionally, we introduce a novel LLM-based evaluation scheme to evaluate VLMs on HiSciVQA which offers deeper insights into model effectiveness surpassing traditional n-gram matching accuracy metrics. We are committed to advancing the field by open-sourcing all code files and the HiSciVQA dataset for the research community.

视觉问答多模态低资源语言微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。