arXiv:2512.15312cs.CLcs.AI2025-12

评测大模型在沸石合成信息抽取中的提示策略效果,发现通用模型难精准提取实验参数。

Evaluating LLMs for Zeolite Synthesis Event Extraction (ZSEE): A Systematic Analysis of Prompting Strategies

  • 对比四种提示策略在沸石合成任务中的表现,聚焦事件分类与参数抽取
  • 事件类型识别准确率80-90%但参数角色和值抽取仅50-65%F1
  • 高级提示策略提升有限,暴露出大模型对专业化学细节理解不足

从沸石合成实验流程中提取结构化信息对材料发现至关重要,但现有方法未系统评估大语言模型(LLMs)在此领域的作用。本文探讨不同提示策略在科学信息抽取任务中的有效性,聚焦四类子任务:事件类型分类(识别合成步骤)、触发词识别(定位事件提及)、参数角色抽取(识别参数类型)和参数值抽取(提取数值)。在包含1,530条标注句子的ZSEE数据集上,评估了六种先进LLM(Gemma-3-12b-it、GPT-5-mini、O4-mini、Claude-Haiku-3.5、DeepSeek推理与非推理版本)及四种提示策略(零样本、少样本、事件特定、反思式)。结果表明,事件类型分类表现良好(F1 80–90%),但细粒度参数抽取表现一般(50–65% F1)。GPT-5-mini在不同提示下性能波动高达11–79%。高级提示策略带来的提升微乎其微,揭示出模型架构的根本局限。错误分析显示存在系统性幻觉、过度泛化及对合成特异性细节捕捉失败的问题。研究证明,尽管大模型具备高层次理解能力,精确提取实验参数仍需领域适配模型,并为科学信息抽取提供量化基准。

原文摘要 · Abstract (English)

Extracting structured information from zeolite synthesis experimental procedures is critical for materials discovery, yet existing methods have not systematically evaluated Large Language Models (LLMs) for this domain-specific task. This work addresses a fundamental question: what is the efficacy of different prompting strategies when applying LLMs to scientific information extraction? We focus on four key subtasks: event type classification (identifying synthesis steps), trigger text identification (locating event mentions), argument role extraction (recognizing parameter types), and argument text extraction (extracting parameter values). We evaluate four prompting strategies - zero-shot, few-shot, event-specific, and reflection-based - across six state-of-the-art LLMs (Gemma-3-12b-it, GPT-5-mini, O4-mini, Claude-Haiku-3.5, DeepSeek reasoning and non-reasoning) using the ZSEE dataset of 1,530 annotated sentences. Results demonstrate strong performance on event type classification (80-90\% F1) but modest performance on fine-grained extraction tasks, particularly argument role and argument text extraction (50-65\% F1). GPT-5-mini exhibits extreme prompt sensitivity with 11-79\% F1 variation. Notably, advanced prompting strategies provide minimal improvements over zero-shot approaches, revealing fundamental architectural limitations. Error analysis identifies systematic hallucination, over-generalization, and inability to capture synthesis-specific nuances. Our findings demonstrate that while LLMs achieve high-level understanding, precise extraction of experimental parameters requires domain-adapted models, providing quantitative benchmarks for scientific information extraction.

信息抽取大模型材料发现提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。