量子模型在经典数据上也能实现高效学习,验证了数据效率的可行性。
Is data-efficient learning feasible with quantum models?
- 构建半人工数据生成工具,定制适合量子核方法的数据集。
- 量子核模型用更少训练样本即可达到与经典模型相当误差。
- 引入谱偏差泛化指标,理论预测与实验结果高度吻合。
在测试量子机器学习(QML)模型时,分析非平凡数据集的重要性日益凸显,但理解数据特征的统一框架仍不明确。本文提出一种数据生成工具,可构造专用于量子核方法(QKMs)的半人工经典数据集。利用该工具,我们发现:在完全经典的数据集上,量子核模型所需的训练样本数可少于经典核方法,且能达到相近的误差水平,为量子模型在经典数据上实现数据高效学习提供了清晰的实证支持。该工具的核心目的是让研究社区能够通过可控实验,探究哪些数据特征特别适合量子模型,只需调整数据生成过程即可。此外,我们将经典核方法中的基于谱偏差的泛化度量引入到QML领域,并证明其预测性能与实际结果高度一致,从而弥合了QML泛化理论与实践之间的关键鸿沟。该工具为系统性探索数据复杂性开辟了道路,有助于深化对量子核模型泛化优势的理解(可扩展至更广泛的QML模型),并推动量子优势的寻找从随机基准测试转向有原则的数据集设计。
原文摘要 · Abstract (English)
The importance of analyzing nontrivial datasets when testing quantum machine learning (QML) models is becoming increasingly prominent in literature, yet a cohesive framework for understanding dataset characteristics remains elusive. In this work, we introduce a data-generation tool that allows to construct semi-artificial classical datasets tailored to quantum kernel methods (QKMs). Using this tool, we show that on fully classical datasets, QKMs can require fewer training examples than classical kernels to reach comparable error, providing clear empirical evidence that data-efficient learning with quantum models is possible on classical data. The main motivation behind this tool is to enable the community to perform controlled studies to figure out which dataset characteristics are particularly fitting for quantum models by tuning the data-generation procedure. Additionally, our study brings a spectral-bias-based generalization metric from classical kernel methods into the QML domain and show that the performance predicted by this metric aligns closely with empirical results, thereby closing an important gap between theory and practice in QML generalization. Our tool paves the way for a systematic exploration of dataset complexities. This could potentially contribute to a deeper understanding of the generalization benefits of QKM models (extendable to a broader family of QML models) and shifts the search for quantum advantage from ad hoc benchmark hunting to principled dataset design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。