数据降维会扭曲量子机器学习模型的真实性能,导致误判。
Influence of Data Dimensionality Reduction Methods on the Effectiveness of Quantum Machine Learning Models
- 对比有无降维,测试多种量子模型和编码方法
- 降维使准确率波动达14%~48%,严重偏差真实表现
- 某些降维方法在特定编码与模型结构下表现更优
数据维度缩减技术常用于量子机器学习(QML)模型中,以应对当前含噪声中等规模量子(NISQ)设备的噪声与比特数限制,以及经典设备对大量量子比特模拟的挑战。然而,该方法在处理大规模数据时适应性差,影响可扩展性。本文分析了不同数据降维方法对多种QML模型的影响,实验覆盖多个生成数据集、量子算法、量子编码方式及降维技术,并评估了准确率、精确率、召回率与F1分数等指标。结果表明,使用降维会导致性能指标严重偏移,从而错误估计实际模型表现。多个因素加剧此问题,包括数据集特性、经典到量子的信息编码方式、特征缩减比例、量子模型中的经典组件及模型结构。实验中,有无降维导致准确率差异范围为14%至48%。此外,某些降维方法在特定编码方式和变分量子线路构造下表现更优。
原文摘要 · Abstract (English)
Data dimensionality reduction techniques are often utilized in the implementation of Quantum Machine Learning models to address two significant issues: the constraints of NISQ quantum devices, which are characterized by noise and a limited number of qubits, and the challenge of simulating a large number of qubits on classical devices. It also raises concerns over the scalability of these approaches, as dimensionality reduction methods are slow to adapt to large datasets. In this article, we analyze how data reduction methods affect different QML models. We conduct this experiment over several generated datasets, quantum machine algorithms, quantum data encoding methods, and data reduction methods. All these models were evaluated on the performance metrics like accuracy, precision, recall, and F1 score. Our findings have led us to conclude that the usage of data dimensionality reduction methods results in skewed performance metric values, which results in wrongly estimating the actual performance of quantum machine learning models. There are several factors, along with data dimensionality reduction methods, that worsen this problem, such as characteristics of the datasets, classical to quantum information embedding methods, percentage of feature reduction, classical components associated with quantum models, and structure of quantum machine learning models. We consistently observed the difference in the accuracy range of 14% to 48% amongst these models, using data reduction and not using it. Apart from this, our observations have shown that some data reduction methods tend to perform better for some specific data embedding methodologies and ansatz constructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。