评测多模态大模型在科学推理任务上的表现,发现Gemini效果最佳。
Scientific Reasoning: Assessment of Multimodal Generative LLMs
- 用ScienceQA数据集评估多模态大模型的科学推理能力。
- Gemini在少上下文时准确率最高,丰富上下文时解释与人类最相似。
- 小模型适配微调无效,用Gemini生成数据训练反而更差。
大型语言模型(LLMs)能够回答问题并处理复杂任务,包括科学领域的问题。我们使用ScienceQA数据集评估了几种多模态大模型(MLLMs)的性能,发现Gemini模型在上下文较少时准确率最高,而在上下文较丰富时,其文本与人类解释的相似度也最高。对小型MLLMs进行适配微调并未带来可靠性能提升。使用Gemini生成的输出进行训练,始终表现不如直接使用原始数据训练。
原文摘要 · Abstract (English)
Large language models (LLMs) can answer questions and reason about complex tasks, also from the scientific domain. We assess several multimodal LLMs (MLLMs) on ScienceQA and find that Gemini models show the highest accuracy with little context, and the highest textual similarity to human explanations with richer context. Adapter-tuning of smaller MLLMs did not lead to any reliable performance. Training from Gemini outputs consistently underperformed training from the original data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。