解决高相关特征空间下的冗余问题,选出有解释力的代表性特征
TangledFeatures: Robust Feature Selection in Highly Correlated Spaces
- 基于特征群中代表性特征识别,降低冗余
- 在丙氨酸二肽数据上提升角度预测性能
- 适合需要可解释性与稳定性分析的研究者
特征选择是模型开发中的关键步骤,影响预测性能和可解释性。然而,大多数常用方法侧重于预测准确率,在存在相关预测变量时性能下降。为此,我们提出 TangledFeatures 框架,用于处理高度相关的特征空间。该方法从纠缠的特征组中识别代表性特征,减少冗余同时保留解释能力。所选特征子集可直接用于下游模型,相比传统方法提供更可解释且稳定的分析基础。我们在丙氨酸二肽数据集上验证了 TangledFeatures 的有效性,应用于骨架扭转角预测,结果表明所选特征对应于具有结构意义的原子内距离,能够解释这些角度的变化。
原文摘要 · Abstract (English)
Feature selection is a fundamental step in model development, shaping both predictive performance and interpretability. Yet, most widely used methods focus on predictive accuracy, and their performance degrades in the presence of correlated predictors. To address this gap, we introduce TangledFeatures, a framework for feature selection in correlated feature spaces. It identifies representative features from groups of entangled predictors, reducing redundancy while retaining explanatory power. The resulting feature subset can be directly applied in downstream models, offering a more interpretable and stable basis for analysis compared to traditional selection techniques. We demonstrate the effectiveness of TangledFeatures on Alanine Dipeptide, applying it to the prediction of backbone torsional angles and show that the selected features correspond to structurally meaningful intra-atomic distances that explain variation in these angles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。