arXiv:2601.19325cs.CVcs.AI2026-01被引 10

用少数据训练出能跨领域推理的科学多模态模型

Innovator-VL: A Multimodal Large Language Model for Scientific Discovery

  • 设计透明可复现的训练流程,从数据到评估全程公开
  • 仅用五百万精选样本即达科学任务竞争力,无需大规模预训练
  • 兼具科学推理与通用视觉能力,适合科研与跨模态研究者

我们提出 Innovator-VL,一个面向科学发现的多模态大语言模型,旨在提升跨学科理解与推理能力,同时保持在通用视觉任务上的优异表现。与依赖海量领域特定预训练和黑箱流程的趋势不同,本工作表明,通过严谨的训练设计与透明的方法论,可在显著降低数据需求的前提下实现强大的科学智能。(i)首先,我们提供一个完全透明、端到端可复现的训练流程,涵盖数据收集、清洗、预处理、监督微调、强化学习及评估,并附详细优化方案,便于社区系统性扩展。(ii)其次,Innovator-VL 展现出卓越的数据效率,在少于五百万条精心筛选样本的情况下,即在多种科学任务上达到竞争力表现,无需大规模预训练。这表明有效推理可通过严谨的数据选择实现,而非盲目扩大规模。(iii)第三,Innovator-VL 在通用视觉、多模态推理和科学基准测试中均表现良好,说明科学对齐可融入统一模型而不损害通用能力。我们的实践表明,即使缺乏大规模数据,也能构建高效、可复现且高性能的科学多模态模型,为未来研究提供切实基础。

原文摘要 · Abstract (English)

We present Innovator-VL, a scientific multimodal large language model designed to advance understanding and reasoning across diverse scientific domains while maintaining excellent performance on general vision tasks. Contrary to the trend of relying on massive domain-specific pretraining and opaque pipelines, our work demonstrates that principled training design and transparent methodology can yield strong scientific intelligence with substantially reduced data requirements. (i) First, we provide a fully transparent, end-to-end reproducible training pipeline, covering data collection, cleaning, preprocessing, supervised fine-tuning, reinforcement learning, and evaluation, along with detailed optimization recipes. This facilitates systematic extension by the community. (ii) Second, Innovator-VL exhibits remarkable data efficiency, achieving competitive performance on various scientific tasks using fewer than five million curated samples without large-scale pretraining. These results highlight that effective reasoning can be achieved through principled data selection rather than indiscriminate scaling. (iii) Third, Innovator-VL demonstrates strong generalization, achieving competitive performance on general vision, multimodal reasoning, and scientific benchmarks. This indicates that scientific alignment can be integrated into a unified model without compromising general-purpose capabilities. Our practices suggest that efficient, reproducible, and high-performing scientific multimodal models can be built even without large-scale data, providing a practical foundation for future research.

多模态模型科学发现数据效率可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。