arXiv:2409.11971cs.CLcond-mat.mtrl-sci2024-09被引 1

用大模型嵌入提取材料属性隐含信息,无需额外训练即可预测性能。

Sampling Latent Material-Property Information From LLM-Derived Embedding Representations

  • 利用大模型文本嵌入挖掘材料文献中的隐含属性信息。
  • 需精准选择上下文线索与对比样本才能有效提取嵌入。
  • 为材料科学提供零样本表示生成新思路,适合跨领域研究者。

从大型语言模型(LLMs)中提取的向量嵌入展现出捕捉文献中隐含信息的潜力。有趣的是,这些嵌入可融入材料表示,有助于基于数据驱动的方法预测材料性能。我们研究了LLM衍生向量在多大程度上能捕获目标信息,以及其在不需额外训练的情况下提供材料属性洞察的可行性。结果表明,尽管LLMs可用于生成反映特定属性信息的表示,但提取有效嵌入需识别最优上下文线索和适当比较对象。尽管存在这一限制,仍显示LLMs在生成有意义的材料科学表示方面具有潜力。

原文摘要 · Abstract (English)

Vector embeddings derived from large language models (LLMs) show promise in capturing latent information from the literature. Interestingly, these can be integrated into material embeddings, potentially useful for data-driven predictions of materials properties. We investigate the extent to which LLM-derived vectors capture the desired information and their potential to provide insights into material properties without additional training. Our findings indicate that, although LLMs can be used to generate representations reflecting certain property information, extracting the embeddings requires identifying the optimal contextual clues and appropriate comparators. Despite this restriction, it appears that LLMs still have the potential to be useful in generating meaningful materials-science representations.

大模型材料科学嵌入表示零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。