首个化学多模态基础模型,让AI懂化学文本、图像、数据等多形式信息。
ChemDFM-X: Towards Large Multimodal Model for Chemistry
- 用计算和模型预测生成多样化化学多模态数据,降低训练成本
- 构建760万条指令微调数据,覆盖多种化学任务与模态
- 实现跨模态理解,为化学通用智能提供关键一步
人工智能工具的快速发展有望为自然科学(包括化学)研究带来前所未有的支持。然而,现有的单模态专用模型或新兴的通用大模型难以涵盖化学领域的广泛数据模态与任务类型。为满足化学家的实际需求,亟需一个跨模态化学通用智能(CGI)系统,作为真正实用的研究助手,发挥大模型的潜力。本文提出首个面向化学的跨模态对话基础模型(ChemDFM-X)。通过近似计算和特定任务模型预测,从初始模态生成多样化的多模态数据,有效构建了大规模训练语料,显著降低开销,形成包含760万条数据的指令微调数据集。经过指令微调后,ChemDFM-X在多种化学任务与数据模态上进行了广泛实验评估,结果表明其具备出色的多模态及跨模态知识理解能力。ChemDFM-X标志着迈向化学所有模态对齐的重要里程碑,是实现化学通用智能的关键一步。
原文摘要 · Abstract (English)
Rapid developments of AI tools are expected to offer unprecedented assistance to the research of natural science including chemistry. However, neither existing unimodal task-specific specialist models nor emerging general large multimodal models (LMM) can cover the wide range of chemical data modality and task categories. To address the real demands of chemists, a cross-modal Chemical General Intelligence (CGI) system, which serves as a truly practical and useful research assistant utilizing the great potential of LMMs, is in great need. In this work, we introduce the first Cross-modal Dialogue Foundation Model for Chemistry (ChemDFM-X). Diverse multimodal data are generated from an initial modality by approximate calculations and task-specific model predictions. This strategy creates sufficient chemical training corpora, while significantly reducing excessive expense, resulting in an instruction-tuning dataset containing 7.6M data. After instruction finetuning, ChemDFM-X is evaluated on extensive experiments of different chemical tasks with various data modalities. The results demonstrate the capacity of ChemDFM-X for multimodal and inter-modal knowledge comprehension. ChemDFM-X marks a significant milestone toward aligning all modalities in chemistry, a step closer to CGI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。