arXiv:2605.05104cond-mat.mtrl-scics.AI2026-05

构建更通用的材料数据集,兼顾当前与未来研究目标。

Building informative materials datasets beyond targeted objectives

论文配图:Building informative materials datasets beyond targeted objectives
图 1 · 摘自论文原文
  • 用多样性感知选择提升材料空间覆盖度。
  • 未使用该框架时,非目标属性性能下降最高达40%。
  • 适合需要长期可用数据的材料发现研究者。

材料科学数据收集成本高昂,数据集的可复用性和长期价值对后续发现至关重要。实践中,研究者常只关注部分性质,忽略其他属性可能导致数据集不适用于未来的学习任务。本文提出一种数据集构建框架,在最大化目标性质信息量的同时,保持对未关注性质的预测性能。该方法采用多样性感知选择,确保材料空间广泛覆盖。在噪声实验数据构建中,不使用该框架时,非目标性质性能相对于随机采样最多下降40%;而应用本框架后,性能提升最高达10%。对于目标性质,无多样性策略时性能相比随机采样最多下降12.5%,而本框架可实现最高25%的提升。将多样性融入数据构建,不仅保留目标性质的信息量,还增强材料覆盖范围,使数据集对已考虑和未考虑的目标均保持高信息量,保障数据质量无偏,缓解后续建模与发现中的冷启动问题。

原文摘要 · Abstract (English)

Materials science data collection can be expensive, making the reuse and long-term utility of datasets critical important for future discovery campaigns. In practice, researchers prioritize a subset of properties due to research interests. However, ignoring a subset of outcomes in data collection campaigns potentially generate datasets poorly suited for future learning tasks. Here, we present a framework for dataset construction that maximizes informativeness for target properties of interest while preserving performance on untargeted ones. Our approach uses diversity-aware selection to ensure broad coverage of the materials space. In noisy experimental dataset construction, we find that without our diversity-aware framework, prediction performance on untargeted properties can degrade by up to 40% relative to random sampling, whereas applying our framework yields improvements of up to 10% . For targeted properties, performance can degrade with respect to random sampling by up to 12.5% without diversity, while our framework achieves gains of up to 25%. Incorporating diversity into dataset construction not only preserves informativeness for the targeted properties, but also improves materials coverage for potential future objectives. As a result, the constructed datasets remain broadly informative across considered and unconsidered outcomes, ensuring unbiased quality entries and mitigating cold-start limitations in subsequent modeling and discovery campaigns.

材料科学数据构建多样性预测性能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。