用量子退火优化筛选错误标签数据,提升训练集质量。
Filtering out mislabeled training instances using black-box optimization and quantum annealing
- 结合黑盒优化与量子退火,迭代评估并过滤错误标签样本。
- 在噪声多数位任务中,有效识别并移除高风险错误标签。
- 物理量子退火器比模拟方法更快更优,适合大规模数据清洗。
本研究提出一种结合代理模型驱动的黑盒优化(BBO)与后处理、量子退火的方法,用于从含噪训练数据集中剔除错误标签样本。真实数据中的错误标签会严重损害模型泛化能力,亟需高效可靠的去噪策略。该方法基于验证损失评估筛选后的训练子集,通过代理模型驱动的BBO迭代优化损失估计,并利用量子退火高效采样低验证误差的多样化训练子集。在噪声多数位任务上的实验表明,该方法能有效优先移除高风险错误标签实例。使用D-Wave的团采样器在物理量子退火器上运行,相比OpenJij的模拟量子退火或Neal的模拟退火,实现更快优化速度和更高品质的训练子集,为提升数据集质量提供了可扩展框架。该方法对监督学习任务有效,未来将拓展至无监督学习、真实数据集及大规模应用。
原文摘要 · Abstract (English)
This study proposes an approach for removing mislabeled instances from contaminated training datasets by combining surrogate model-based black-box optimization (BBO) with postprocessing and quantum annealing. Mislabeled training instances, a common issue in real-world datasets, often degrade model generalization, necessitating robust and efficient noise-removal strategies. The proposed method evaluates filtered training subsets based on validation loss, iteratively refines loss estimates through surrogate model-based BBO with postprocessing, and leverages quantum annealing to efficiently sample diverse training subsets with low validation error. Experiments on a noisy majority bit task demonstrate the method's ability to prioritize the removal of high-risk mislabeled instances. Integrating D-Wave's clique sampler running on a physical quantum annealer achieves faster optimization and higher-quality training subsets compared to OpenJij's simulated quantum annealing sampler or Neal's simulated annealing sampler, offering a scalable framework for enhancing dataset quality. This work highlights the effectiveness of the proposed method for supervised learning tasks, with future directions including its application to unsupervised learning, real-world datasets, and large-scale implementations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。