用图文数据提升材料科学模型的推理能力,实测效果显著。
MATRIX: A Multimodal Benchmark and Post-Training Framework for Materials Science
- 用图文配对数据进行后训练,强化模型对实验图像的理解
- 图文联合训练使实验解读准确率提升10%-25%,文本推理提5%-16%
- 适合需要跨模态理解的材料科学与AI研究者使用
材料科学中的科学推理需融合多模态实验证据与物理理论。现有基准难以评估在后训练中加入视觉实验数据是否能提升机制性解释推理能力。我们提出MATRIX,一个用于材料科学推理的多模态基准,涵盖基础理论、研究级推理及多种表征模态下的真实实验特征解读。通过该基准,我们对比仅使用结构化文本与同时引入配对实验图像的后训练效果。尽管多模态数据量较小,但视觉监督使实验解读准确率提升10%-25%,在纯文本科学推理任务上获得5%-16%的增益。结果表明,这些提升依赖于后训练中正确的图文对齐,凸显跨模态表征迁移的重要性。此外,在ScienceQA和PubMedQA上也观察到一致提升,说明结构化多模态后训练的优势可推广至更广泛领域。数据集与模型已公开于Hugging Face。
原文摘要 · Abstract (English)
Scientific reasoning in materials science requires integrating multimodal experimental evidence with underlying physical theory. Existing benchmarks make it difficult to assess whether incorporating visual experimental data during post-training improves mechanism-grounded explanation reasoning beyond text-only supervision. We introduce MATRIX, a multimodal benchmark for materials science reasoning that evaluates foundational theory, research-level reasoning, and the interpretation of real experimental artifacts across multiple characterization modalities. Using MATRIX as a controlled diagnostic, we isolate the effect of visual grounding by comparing post-training on structured materials science text alone with post-training that incorporates paired experimental images. Despite using relatively small amounts of multimodal data, visual supervision improves experimental interpretation by 10-25% and yields 5-16% gains on text-only scientific reasoning tasks. Our results demonstrate that these improvements rely on correct image-text alignment during post-training, highlighting cross-modal representational transfer. We also observe consistent improvements on ScienceQA and PubMedQA, demonstrating that the benefits of structured multimodal post-training extend beyond materials science. The MATRIX dataset is available at https://huggingface.co/datasets/radical-ai/MATRIX and the model at https://huggingface.co/radical-ai/MATRIX-PT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。