arXiv:2512.01509quant-phcs.LG2025-12被引 4

用自编码器压缩高维粒子物理数据,提升量子分类器性能

Learning Reduced Representations for Quantum Classifiers

  • 用六种传统方法和五种自编码器降维,保留关键特征
  • 自编码器比传统方法提升40%分类准确率,新设计的Sinkclass表现最佳
  • 为高维数据量子机器学习提供可复用的降维方案,适合物理与数据科学者

当前量子机器学习算法难以处理高维特征数据集。一个直接解决方案是在输入量子算法前使用降维方法。本文在包含67个特征的粒子物理数据集上,测试了六种传统特征提取算法和五种基于自编码器的降维模型,并将其用于训练量子支持向量机,解决大型强子对撞机中质子碰撞是否产生希格斯玻色子的二分类问题。结果表明,自编码器能学习到更优的低维表示,所提出的Sinkclass自编码器相比基线方法性能提升40%。本研究拓展了量子机器学习在更多数据集上的适用性,并提供了有效的降维实践指南。

原文摘要 · Abstract (English)

Data sets that are specified by a large number of features are currently outside the area of applicability for quantum machine learning algorithms. An immediate solution to this impasse is the application of dimensionality reduction methods before passing the data to the quantum algorithm. We investigate six conventional feature extraction algorithms and five autoencoder-based dimensionality reduction models to a particle physics data set with 67 features. The reduced representations generated by these models are then used to train a quantum support vector machine for solving a binary classification problem: whether a Higgs boson is produced in proton collisions at the LHC. We show that the autoencoder methods learn a better lower-dimensional representation of the data, with the method we design, the Sinkclass autoencoder, performing 40% better than the baseline. The methods developed here open up the applicability of quantum machine learning to a larger array of data sets. Moreover, we provide a recipe for effective dimensionality reduction in this context.

量子机器学习降维粒子物理自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。