多模型验证发现,电纺丝中只有溶液浓度影响稳定,其他参数重要性依赖模型。
Cross-Model Consistency of Feature Importance in Electrospinning: Separating Robust from Model-Dependent Features

- 用21种模型+SHAP统一评估特征重要性
- 浓度重要性一致(变异度0),流速和电压差异大(变异度>0.9)
- 提醒仅靠单一模型解读风险高,适合材料工艺研究者
电纺丝是高度敏感的制造过程,操作参数微小变化会显著影响纤维形态与材料性能。机器学习(ML)被广泛用于建模工艺-结构关系并识别参数相对重要性。然而,多数研究依赖单一模型,隐含假设特征重要性具鲁棒性。本研究基于96组聚乙烯醇(PVA)电纺实验的定制数据集,系统评估了21种代表线性、树基、核方法、神经网络和实例基等模型家族的特征重要性一致性。所有模型均使用SHAP值统一计算特征重要性,并通过秩统计分析量化跨模型一致性。结果表明,预测性能与解释可靠性是根本不同的属性:尽管多个模型预测精度相当,但特征重要性排序存在显著差异。溶液浓度为最稳健且一致影响的参数(变异度=0),而流速和施加电压排名变异度>0.9,表明其重要性强烈依赖模型。这些发现说明,单一模型得出的特征重要性可能不可靠,尤其在小样本实验数据下,强调跨模型验证对实现可信解释的重要性。
原文摘要 · Abstract (English)
Electrospinning is a highly sensitive fabrication process in which small variations in operating parameters can significantly influence fiber morphology and material performance. Machine learning (ML) methods are increasingly employed to model these process-structure relationships and to identify the relative importance of processing variables. However, most existing studies rely on a single ML model, implicitly assuming that the resulting feature importance is robust and reproducible. In this study, the consistency of feature importance across multiple ML model families was systematically evaluated using a curated dataset of 96 polyvinyl alcohol (PVA) electrospinning experiments. Twenty-one ML models representing linear, tree-based, kernel-based, neural network, and instance-based approaches were trained and compared. To provide a unified interpretability framework, SHAP (SHapley Additive exPlanations) values were used to calculate feature importance consistently across all models. A rank-based statistical analysis was then performed to quantify inter-model agreement and assess the robustness of parameter rankings. The results demonstrate that predictive performance and interpretive reliability are fundamentally distinct properties. Although several models achieved comparable predictive accuracy, substantial differences were observed in their feature importance rankings. Solution concentration emerged as the most robust and consistently influential parameter (variability = 0), whereas flow rate and applied voltage exhibited high ranking variability (variability > 0.9), indicating strong model dependence. These findings suggest that feature importance derived from a single ML model may be unreliable, particularly for small experimental datasets, and highlight the importance of cross-model validation for achieving trustworthy interpretation in ML-assisted electrospinning research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。