arXiv:2510.00795cs.AI2025-10中稿 · The AI for Acceler…被引 2

ChemX为化学信息提取提供10个专家验证数据集,推动智能体系统精准处理分子与纳米材料数据。

Benchmarking Agentic Systems in Automated Scientific Information Extraction with ChemX

  • 构建10个领域专家验证的化学数据集,聚焦纳米材料与小分子信息提取。
  • 发现现有智能体在术语、表格和上下文歧义处理上仍存显著缺陷。
  • 适合从事化学信息自动化、智能体评估及领域模型优化的研究者参考。

智能体系统在人工智能中取得重要进展,广泛应用于自动化数据提取。然而,由于化学数据固有的异质性,化学信息提取仍是重大挑战。当前通用及领域专用的智能体方法在该领域表现有限。为此,我们提出ChemX,一个包含10个手工构建且经领域专家验证的数据集,专注于纳米材料与小分子。这些数据集旨在严格评估并提升化学领域的自动化提取方法。为展示其价值,我们开展大规模基准测试,对比ChatGPT Agent等前沿智能体系统,以及专用于化学数据提取的智能体。此外,我们提出一种单智能体方法,可在提取前精细控制文档预处理。还评估了GPT-5与GPT-5 Thinking等现代基线模型的性能,以比较其与智能体方法的差异。实证结果表明,化学信息提取仍面临持续挑战,尤其在处理领域术语、复杂表格与示意图,以及上下文依赖的歧义问题上。ChemX基准为推进化学领域自动化信息提取提供了关键资源,挑战现有方法的泛化能力,并为有效评估策略提供深刻洞见。

原文摘要 · Abstract (English)

The emergence of agent-based systems represents a significant advancement in artificial intelligence, with growing applications in automated data extraction. However, chemical information extraction remains a formidable challenge due to the inherent heterogeneity of chemical data. Current agent-based approaches, both general-purpose and domain-specific, exhibit limited performance in this domain. To address this gap, we present ChemX, a comprehensive collection of 10 manually curated and domain-expert-validated datasets focusing on nanomaterials and small molecules. These datasets are designed to rigorously evaluate and enhance automated extraction methodologies in chemistry. To demonstrate their utility, we conduct an extensive benchmarking study comparing existing state-of-the-art agentic systems such as ChatGPT Agent and chemical-specific data extraction agents. Additionally, we introduce our own single-agent approach that enables precise control over document preprocessing prior to extraction. We further evaluate the performance of modern baselines, such as GPT-5 and GPT-5 Thinking, to compare their capabilities with agentic approaches. Our empirical findings reveal persistent challenges in chemical information extraction, particularly in processing domain-specific terminology, complex tabular and schematic representations, and context-dependent ambiguities. The ChemX benchmark serves as a critical resource for advancing automated information extraction in chemistry, challenging the generalization capabilities of existing methods, and providing valuable insights into effective evaluation strategies.

化学信息提取智能体系统数据集基准纳米材料

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。