首次系统测量材料生成模型的规模外推极限,揭示不同模型的失效规律。
How Far Can You Grow? Characterizing the Extrapolation Frontier of Graph Generative Models for Materials Science
- 构建连续尺度的纳米颗粒基准数据集,分离几何外推影响
- 发现模型在超出训练尺度后全局误差上升约13%,局部键长误差可增2倍
- 提出多维失效分析方法,帮助预测模型可扩展性边界
所有晶体材料生成模型均存在一个关键结构尺寸上限,超过该限值生成结果便不可靠,我们称之为外推前沿。尽管这对纳米材料设计至关重要,但该前沿从未被系统测量。本文提出RADII,一个包含约7.5万种晶体制纳米颗粒结构(33-11,298原子)的半径分辨基准,将半径作为连续缩放变量,在无泄漏划分下追踪生成质量从分布内到分布外的变化。每个模型基于目标组成与原子数进行条件生成,隔离几何外推作为唯一评估变量。RADII提供前沿特异性诊断:每半径误差曲线定位各架构的缩放上限,表面-体相分解区分边界与体部失败,跨指标序列揭示结构保真度最先崩溃的部分。对五种先进架构的评测显示:(i) 表现良好的模型在训练半径外全局位置误差增加约13%,而发散模型在全尺度上保真度差,局部键长误差从几乎无退化到增长超2倍;(ii) 无两个架构共享失效顺序,表明前沿是受模型家族决定的多维曲面;(iii) 表现良好模型遵循预期几何缩放指数α≈1/3,其分布内拟合可预测分布外误差,使前沿具备可预测性。将MatterGen扩展至发布参数量虽稳定采样,但未消除前沿;而DiffCSP在发布规模下仍不稳定。这些发现确立输出规模为几何生成模型的一级评估维度。代码与数据:https://github.com/KurbanIntelligenceLab/RADII。
原文摘要 · Abstract (English)
Every generative model for crystalline materials harbors a critical structure size beyond which its outputs become unreliable; we call this the extrapolation frontier. Despite its consequences for nanomaterial design, this frontier has never been systematically measured. We introduce RADII, a radius-resolved benchmark of ~75,000 crystal-derived nanoparticle structures (33-11,298 atoms) that treats radius as a continuous scaling knob, tracing generation quality from in- to out-of-distribution under leakage-free splits. Each model is conditioned on target composition and atom count, isolating geometric extrapolation as the evaluation variable. RADII provides frontier-specific diagnostics: per-radius error profiles pinpoint each architecture's scaling ceiling, surface-interior decomposition separates boundary from bulk failures, and cross-metric sequencing reveals which aspect of structural fidelity breaks first. Benchmarking five state-of-the-art architectures, we find that: (i) well-behaved models degrade by ~13% in global positional error beyond training radii, while divergent models show poor fidelity across scales, with local bond fidelity ranging from negligible degradation to over 2x error growth; (ii) no two architectures share a failure sequence, revealing the frontier as a multi-dimensional surface shaped by model family; and (iii) well-behaved models follow the expected geometric scaling exponent alpha ~ 1/3, whose in-distribution fit predicts out-of-distribution error, making frontiers forecastable. Scaling MatterGen to its published parameter count stabilizes sampling but does not close the frontier, while DiffCSP remains unstable at published scale. These findings establish output scale as a first-class evaluation axis for geometric generative models. Code and data: https://github.com/KurbanIntelligenceLab/RADII.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。