构建科学领域多模态指令遵循评测基准,揭示模型短板。
SciMIF: Understanding Multimodal Instruction Following in Scientific Domains

- 基于10类约束构建科学任务分类体系,系统化注入指令。
- 发现化学任务中模型表现最差,且增大模型规模无助于提升约束遵守率。
- 适用于评估科学场景下多模态大模型的推理与指令理解能力。
理解科学领域中的指令遵循能力对有效利用多模态大语言模型(MLLMs)推动科学进步至关重要。本文提出SciMIF,一个新型基准,用于评估MLLMs在复杂科学指令下的遵循能力。基于对5个代表性科学领域中22项任务的分析,我们构建了包含10类约束的综合分类体系,涵盖通用功能需求与学科特异性特征。据此,设计高保真指令注入流程,系统扩充现有科学数据集。在多个主流闭源与开源MLLM上开展全面实验。结果表明,不同科学领域间性能差异显著,化学任务对当前模型挑战最大;增加模型规模并未带来约束遵守率的相应提升,现有模型仍严重难以处理细粒度约束及需深度学科知识的应用。SciMIF填补了科学领域多模态指令遵循评估的空白,为未来MLLM在严谨科学应用中的改进奠定基础。数据与代码将公开于https://github.com/shenye7436/SciMIF。
原文摘要 · Abstract (English)
Understanding instruction-following capabilities in scientific domains is essential for effectively leveraging Multimodal Large Language Models (MLLMs) to advance the development of scientific fields. In this work, we introduce SciMIF, a novel benchmark designed to evaluate the capability of MLLMs in following complex scientific instructions. Specifically, based on an extensive analysis of 22 distinct tasks across 5 representative scientific disciplines, we propose a comprehensive taxonomy comprising 10 constraint groups that captures both general functional requirements and discipline-specific characteristics. Guided by this taxonomy, we develop a high-fidelity instruction injection pipeline to systematically augment existing scientific datasets. We conduct comprehensive experiments on multiple state-of-the-art closed-source and open-source MLLMs. Our findings reveal significant performance disparities across different scientific disciplines, with chemistry posing greater challenges for current MLLMs. Furthermore, we observe that increasing the model scale does not yield corresponding improvements in constraint adherence, and current models still struggle severely with fine-grained constraints and instructions requiring the deep application of disciplinary knowledge. SciMIF fills the current void in evaluating multimodal instruction adherence within scientific domains, laying a crucial foundation for future enhancements of MLLMs in rigorous scientific applications. Data and code will be released at https://github.com/shenye7436/SciMIF .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。