让文字描述直接检索晶体结构,打通材料文本与结构的跨模态连接
Bridging Text and Crystal Structures: Literature-driven Contrastive Learning for Materials Science
- 用文本和晶体结构对比学习,构建可语义检索的材料嵌入空间
- 基于40万+文献数据训练,实现以文搜材的精准匹配
- 适合材料科研人员快速查找功能相似的晶体结构
理解结构-性能关系是材料发现与开发中的核心挑战。近年来,材料信息学尝试通过晶体结构的隐式嵌入空间捕捉其性质与功能上的相似性,但抽象的特征嵌入难以直观探索浩瀚的材料空间。本文提出对比语言-结构预训练(CLaSP),建立晶体结构与文本间的跨模态嵌入空间,旨在实现:1)捕捉晶体结构在性质与功能上的相似性;2)支持用户以自然语言描述作为查询,直观检索材料。为弥补缺乏结构-文本配对数据的问题,CLaSP利用包含超过40万条晶体结构及其论文标题与摘要的公开文献数据集进行训练。通过基于文本的晶体结构筛选和嵌入空间可视化,验证了CLaSP的有效性。
原文摘要 · Abstract (English)
Understanding structure-property relationships is an essential yet challenging aspect of materials discovery and development. To facilitate this process, recent studies in materials informatics have sought latent embedding spaces of crystal structures to capture their similarities based on properties and functionalities. However, abstract feature-based embedding spaces are human-unfriendly and prevent intuitive and efficient exploration of the vast materials space. Here we introduce Contrastive Language--Structure Pre-training (CLaSP), a learning paradigm for constructing crossmodal embedding spaces between crystal structures and texts. CLaSP aims to achieve material embeddings that 1) capture property- and functionality-related similarities between crystal structures and 2) allow intuitive retrieval of materials via user-provided description texts as queries. To compensate for the lack of sufficient datasets linking crystal structures with textual descriptions, CLaSP leverages a dataset of over 400,000 published crystal structures and corresponding publication records, including paper titles and abstracts, for training. We demonstrate the effectiveness of CLaSP through text-based crystal structure screening and embedding space visualization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。