arXiv:2603.08881cond-mat.mtrl-scics.CL2026-03

用文本嵌入筛选复杂电催化剂,无需实验数据也能高效缩小候选范围。

From Word2Vec to Transformers: Text-Derived Composition Embeddings for Filtering Combinatorial Electrocatalysts

  • 用科学文本生成元素嵌入,通过线性组合构建材料组合嵌入。
  • 在15个材料库中,轻量Word2Vec模型实现最高候选数削减,性能接近最优。
  • 适合无标签条件下快速筛选新型电催化材料的科研人员使用。

成分复杂的固溶体电催化剂覆盖广阔的成分空间,单一材料体系的候选组合数量可能远超可实验测量的范围。本文评估了一种无标签筛选策略:利用科学文本训练的嵌入表示每个成分,并根据与两个性能概念(导电性、介电性)的相似度优先排序候选。对比了基于语料库训练的Word2Vec基线与基于Transformer的嵌入方法,其中成分通过元素嵌入的线性组合或短提示词编码。与‘概念方向’的相似性构成二维描述符空间,采用对称帕累托前沿选择法筛选候选子集,无需电化学标签。在包含贵金属合金和多组分氧化物在内的15个材料库上进行评估。结果表明,轻量级的Word2Vec基线模型(采用元素嵌入线性组合)通常能实现最多的候选组合缩减,同时保持接近最佳实测性能。

原文摘要 · Abstract (English)

Compositionally complex solid solution electrocatalysts span vast composition spaces, and even one materials system can contain more candidate compositions than can be measured exhaustively. Here we evaluate a label-free screening strategy that represents each composition using embeddings derived from scientific texts and prioritizes candidates based on similarity to two property concepts. We compare a corpus-trained Word2Vec baseline with transformer-based embeddings, where compositions are encoded either by linear element-wise mixing or by short composition prompts. Similarities to `concept directions', the terms conductivity and dielectric, define a 2-dimensional descriptor space, and a symmetric Pareto-front selection is used to filter candidate subsets without using electrochemical labels. Performance is assessed on 15 materials libraries including noble metal alloys and multicomponent oxides. In this setting, the lightweight Word2Vec baseline, which uses a simple linear combination of element embeddings, often achieves the highest number of reductions of possible candidate compositions while staying close to the best measured performance.

电催化文本嵌入无监督筛选材料发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。