arXiv:2609.01621cs.IRcond-mat.mtrl-sci2026-09

文献数据误导AI材料发现,因隐藏错误难以察觉

When Literature Data Mislead Artificial Intelligence in Materials Discovery

  • 通过追踪文献数据源,发现文本与图表、单位等存在多类不一致
  • 模糊报告可导致导电率误差达100倍,影响模型训练可靠性
  • 呼吁建立可追溯的数据标注与验证机制,保障AI科研可信度

人工智能(AI)越来越多地将科学文献作为构建数据库、训练预测模型和指导发现的数据来源。然而,文献衍生数据集通常假设报告的实验值在内部一致且可直接复用。本文以固态电解质导电率数据为例,分析这一假设的合理性。通过追踪原始论文到整理数据集的过程,识别出重复出现的图文不符、坐标轴标注模糊、单位不一致以及测量背景缺失等问题。这些差异虽数值上合理,难以通过常规预处理发现,但在数据库构建和机器学习应用中会作为结构化标签噪声传播。跨数据库案例显示,模糊报告可导致导电率误差高达100倍。本研究将数据准确性重新定位为人工智能驱动发现的基础性要求,并倡导可追溯的报告、整理与验证实践,以实现可复用的科学数据。

原文摘要 · Abstract (English)

Artificial intelligence (AI) increasingly treats scientific literature as a data source for building databases, training predictive models, and guiding discovery. Yet literature-derived datasets often assume that reported experimental values are internally consistent and directly reusable. Here, we analyze this assumption using solid electrolyte (SE) conductivity data as a representative materials-science case. By tracing values from source articles to curated datasets, we identify recurrent text-figure mismatches, ambiguous axis annotations, unit inconsistencies, and missing measurement context. These discrepancies are often numerically plausible and therefore difficult to detect through routine preprocessing, but they can propagate as structured label noise during database construction and machine-learning reuse. A cross-database example shows how ambiguous reporting can create a 100-fold conductivity error. Our analysis reframes data accuracy as an infrastructure requirement for artificial-intelligence-driven discovery and motivates traceable reporting, curation, and validation practices for reusable scientific data. Keywords: AI for science; Data reliability; Scientific databases; Structured label noise; Literature-derived data; Materials informatics; Solid electrolytes

AI科研数据可靠性材料信息学文献数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。