arXiv:2606.05139cs.LG2026-06

首个生物数据无监督表征学习超参优化基准,助力算法高效研发

BBOmix: A Tabular Benchmark for Hyperparameter Optimization of Unsupervised Biological Representation Learning

  • 构建包含10.5万次评估的开源基准,覆盖4种自编码器与7类组学数据
  • 发现重建损失与下游任务性能相关性较弱,揭示现有优化目标缺陷
  • 支持多保真度与迁移学习方法对比,适合生物信息与自动化机器学习研究者

高通量测序技术推动了大规模、高维度组学数据的发展。深度无监督学习架构,尤其是自编码器(AEs),在该领域广泛用于降维和表征学习。然而,自编码器对结构设计和超参数极为敏感,且无监督优化通常依赖重建损失,该指标可能无法有效反映下游任务表现。全面的超参数优化(HPO)计算成本高昂,导致研究者常采用次优默认配置。为促进大规模无监督HPO研究的可及性,我们提出BBOmix——首个针对真实生物数据的无监督表征学习开源表格基准。该基准涵盖来自TCGA和SCHC数据集的四种自编码器架构与七种多组学模态,共105,000次评估。我们量化了重建损失与下游任务性能之间的相关性,并系统评估了当前先进的单保真度、多保真度及迁移学习HPO方法,为无监督生物表征学习研究建立了严格基线。

原文摘要 · Abstract (English)

The rapid advancement of high-throughput sequencing has led to large, high-dimensional omics datasets. Deep unsupervised learning architectures, particularly Autoencoders (AEs), are increasingly used for dimensionality reduction and representation learning in this domain. However, AEs are highly sensitive to architectural choices and hyperparameters, and unsupervised optimization typically relies on reconstruction loss, which may be a poor proxy for downstream utility. Exhaustive hyperparameter optimization (HPO) is computationally expensive, leading researchers to frequently rely on suboptimal default configurations. To democratize access to large-scale unsupervised HPO research, we introduce $\textbf{BBOmix}$, the first open-source tabular benchmark for unsupervised representation learning on real-world biological data. Our benchmark includes 105,000 evaluations across four AE architectures and seven multi-omics modalities from the TCGA and SCHC datasets. We quantify the correlation between reconstruction loss and downstream task performance and provide an extensive evaluation of state-of-the-art single-fidelity, multi-fidelity, and transfer learning HPO methods, establishing a rigorous baseline for future research in unsupervised biological representation learning.

超参优化生物信息自编码器多组学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。