用量子电路生成合成数据,提升不平衡表格数据的分类效果。
Quantum-Enhanced Synthetic Data Generation Using Quantum Circuit Born Machines for Imbalanced Tabular Learning
- 用量子电路建模复杂分布,生成更贴近真实数据的合成样本。
- 合成数据使少数类召回率提升10%-25%,F1分数提高5%-15%。
- 适合处理低维结构化表格数据的不平衡问题,可作传统方法补充。
数据稀缺与类别不平衡是机器学习中持续存在的挑战,会降低模型泛化能力并引入预测偏差。本文提出一种混合量子-经典框架,利用量子电路玻恩机(QCBM)生成合成数据以应对这些局限。该方法在参数化变分量子电路中利用量子叠加与纠缠特性,建模经典生成方法难以捕捉的复杂概率分布。实验基于鸢尾花和电信客户流失两个表格基准数据集展开。预处理包括归一化与基于PCA的降维,以实现量子电路的高效基编码。通过梯度下降的参数移位优化规则最小化真实数据与生成数据分布间的KL散度来训练QCBM。在少数类占比40%-50%时,用合成样本扩充训练数据,使F1分数提升约5%-15%,少数类召回率提升10%-25%。跨域评估(用合成数据训练、真实数据测试;反之亦然)显示性能差距仅3%-10%,表明分布保真度高。与经典过采样方法SMOTE、Borderline-SMOTE、KMeansSMOTE及SVM-SMOTE对比,在电信数据集上取得相当的分类性能,并呈现更低的最大均值差异(MMD),说明在某些不平衡场景下具有更优的结构相似性。研究结果表明,QCBM可作为数据增强的有效互补工具,尤其适用于低维结构化表格数据的类别不平衡问题。
原文摘要 · Abstract (English)
Data scarcity and class imbalance are persistent challenges in machine learning that degrade model generalization and introduce predictive bias. We present a hybrid quantum-classical framework for synthetic data generation using a Quantum Circuit Born Machine (QCBM) to address these limitations. The proposed approach exploits quantum mechanical properties -- superposition and entanglement -- within a parameterized variational quantum circuit to model complex probability distributions that are difficult for classical generative methods to capture. Experiments are conducted on two tabular benchmark datasets: the Iris dataset and the Telco Customer Churn dataset. Preprocessing includes normalization and PCA-based dimensionality reduction to enable efficient basis encoding for quantum circuits. The QCBM is trained by minimizing Kullback-Leibler (KL) divergence between real and generated data distributions using a gradient-based parameter-shift optimization rule. Augmenting training data with QCBM-generated synthetic samples at 40-50% of the minority class improves F1-score by approximately 5-15% and minority-class recall by 10-25%. Cross-domain evaluations (Train on Synthetic, Test on Real; and Train on Real, Test on Synthetic) reveal a performance gap of only 3-10%, indicating strong distributional fidelity. Comparative analysis against classical oversampling methods -- SMOTE, Borderline-SMOTE, KMeansSMOTE, and SVM-SMOTE -- shows that QCBM achieves competitive classification performance and produces lower Maximum Mean Discrepancy (MMD) on the Telco dataset, suggesting superior structural similarity in certain imbalanced settings. These findings establish QCBM as a viable complementary tool for data augmentation, particularly for low-dimensional structured tabular data with class imbalance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。