用量子计算优化少数类数据生成,提升不平衡数据集的分类效果
Quantum SMOTE with Angular Outliers: Redefining Minority Class Handling
- 基于量子旋转与角度分布生成新样本,避免传统聚类方法
- 在30%-36%合成率下性能超越原方法50%合成率的表现
- 适合处理高维数据且对边缘案例识别有显著提升
本文提出 Quantum-SMOTEV2,一种无需 K-Means 聚类的量子机器学习数据增强方法,通过交换测试与围绕单一中心点的量子旋转生成样本,聚焦少数类数据点的角度分布及角异常值(AOL)。实验表明,在中等合成比例(30%-36%)下模型性能显著提升,此前需高达50%合成率才能达到同等效果。该方法保留前代版本的旋转角度、少数类占比和分割因子等核心参数,支持针对特定数据集定制化调整。其采用紧凑的交换测试与低深度量子电路,具备良好的可扩展性。在公开的 Cell-to-Cell Telecom 数据集上,结合随机森林(RF)、K-近邻(KNN)与神经网络(NN)评估显示,引入角异常值能小幅但稳定提升准确率、F1 分数、AUC-ROC 与 AUC-PR,验证了 Quantum-SMOTEV2 在边缘案例处理上的有效性。
原文摘要 · Abstract (English)
This paper introduces Quantum-SMOTEV2, an advanced variant of the Quantum-SMOTE method, leveraging quantum computing to address class imbalance in machine learning datasets without K-Means clustering. Quantum-SMOTEV2 synthesizes data samples using swap tests and quantum rotation centered around a single data centroid, concentrating on the angular distribution of minority data points and the concept of angular outliers (AOL). Experimental results show significant enhancements in model performance metrics at moderate SMOTE levels (30-36%), which previously required up to 50% with the original method. Quantum-SMOTEV2 maintains essential features of its predecessor (arXiv:2402.17398), such as rotation angle, minority percentage, and splitting factor, allowing for tailored adaptation to specific dataset needs. The method is scalable, utilizing compact swap tests and low depth quantum circuits to accommodate a large number of features. Evaluation on the public Cell-to-Cell Telecom dataset with Random Forest (RF), K-Nearest Neighbours (KNN) Classifier, and Neural Network (NN) illustrates that integrating Angular Outliers modestly boosts classification metrics like accuracy, F1 Score, AUC-ROC, and AUC-PR across different proportions of synthetic data, highlighting the effectiveness of Quantum-SMOTEV2 in enhancing model performance for edge cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。