arXiv:2502.08785cs.LGcs.SD2025-02

用进化算法优化听力损失筛查数据,提升模型性能。

Decision Tree Based Wrappers for Hearing Loss

  • 以决策树为代理,用进化算法自动筛选和构造特征。
  • 仅用57个特征达到76.2%的平衡准确率,单特征也能达72.8%。
  • 适合医疗筛查中数据维度高、需可解释性的场景。

听力学机构正使用机器学习模型识别高风险人群。特征工程(FE)通过优化数据提升模型效果,其中进化方法在特征选择与构造中表现优异。本文基于决策树构建进化特征工程框架(FEDORA),应用于听力损失(HL)数据集,成功降低数据维度并保持基线性能。相比传统方法,FEDORA在使用57个特征时实现最高76.2%的平衡准确率;同时生成的个体仅用一个特征即达到72.8%的平衡准确率。

原文摘要 · Abstract (English)

Audiology entities are using Machine Learning (ML) models to guide their screening towards people at risk. Feature Engineering (FE) focuses on optimizing data for ML models, with evolutionary methods being effective in feature selection and construction tasks. This work aims to benchmark an evolutionary FE wrapper, using models based on decision trees as proxies. The FEDORA framework is applied to a Hearing Loss (HL) dataset, being able to reduce data dimensionality and statistically maintain baseline performance. Compared to traditional methods, FEDORA demonstrates superior performance, with a maximum balanced accuracy of 76.2%, using 57 features. The framework also generated an individual that achieved 72.8% balanced accuracy using a single feature.

听力损失特征工程进化算法决策树

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。