arXiv:2605.20413cs.LG2026-05

通过监督重构降维,提升小样本植物表型数据的分类性能

Supervised Latent Restructuring for Small-Data Quantum Learning in Plant Phenomics

论文配图:Supervised Latent Restructuring for Small-Data Quantum Learning in Plant Phenomics
图 1 · 摘自论文原文
  • 先用PCA压缩到64维,再用LDA构建11维有监督潜在空间
  • 重构后轮廓系数从-0.006提升至0.197,几何可分性显著改善
  • 适合小样本生物图像分类,尤其关注量子学习的表征设计

高维生物数据常面临特征维度远超样本数量的问题,导致极小样本情形下分类困难。此时核方法在潜在压缩无法保留类别区分结构时会丧失判别力。本文研究细粒度植物表型问题,提出混合流程:将1280维深度图像嵌入压缩至64维PCA空间,再通过线性判别分析(LDA)重构为11维有监督潜在空间,并在NVIDIA L40S硬件上实现GPU加速的量子核对齐(QKA)。实验表明,重构后压缩表示的几何可分性显著提升,轮廓系数从原始嵌入空间的0.003、PCA-64空间的-0.006增至监督LDA-11空间的0.197。然而下游经典评估显示存在压缩权衡:线性SVM与XGBoost在重构空间中表现提升,而RBF-SVM与随机森林则在相同11维瓶颈下退化。在受限优化预算下,该场景中的量子核对齐仍具挑战,表明仅靠潜在几何不足以实现强可训练量子性能。研究将表示几何置于小样本量子学习的核心设计变量位置,揭示从高度压缩的生物表征中恢复非线性判别结构的现实困难。

原文摘要 · Abstract (English)

High-dimensional biological data often exhibit a severe mismatch between feature dimensionality and sample size, making reliable classification difficult in extremely small-data regimes. In these settings, kernel methods can lose discriminative power when latent compression fails to preserve class-separating structure. We study this problem in fine-grained plant phenomics and propose a hybrid workflow that compresses 1280-dimensional deep image embeddings into a 64-dimensional PCA space and then restructures them into an 11-dimensional supervised latent space using Linear Discriminant Analysis (LDA), followed by GPU-accelerated Quantum Kernel Alignment (QKA) on NVIDIA L40S hardware. Empirically, supervised latent restructuring substantially improves the geometric separability of the compressed representation, increasing the Silhouette coefficient from 0.003 in the raw embedding space and -0.006 in PCA-64 to 0.197 in the supervised LDA-11 space. However, downstream classical evaluation reveals a clear compression trade-off: Linear SVM and XGBoost improve in the restructured latent space, whereas RBF-SVM and Random Forest degrade under the same 11-dimensional bottleneck. Under a constrained optimization budget, QKA in this regime remains challenging, indicating that latent geometry alone is not sufficient for strong trainable quantum performance. These findings position representation geometry as a central design variable in small-data quantum learning and expose the practical difficulty of recovering nonlinear discriminative structure from aggressively compressed biological representations.

小样本学习量子核对齐表型分析降维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。