BALDUR用贝叶斯方法融合多模态生物数据,在小样本下实现可解释分类。
Unified Bayesian representation for high-dimensional multi-modal biomedical data for small-sample classification
- 在统一潜在空间中融合多模态数据,自动筛选关键特征
- 双核机制应对小样本高维问题,性能优于现有模型
- 线性结构确保结果可解释,适合生物标志物发现
我们提出BALDUR,一种新型贝叶斯算法,用于处理高维多模态生物医学数据的小样本分类问题,并提供可解释的解决方案。该模型将不同数据视图映射到共同潜在空间,提取对分类任务相关的信息,同时剔除无关或冗余特征/数据视图。为在小样本场景中实现泛化能力,BALDUR高效集成双核机制于各数据视图,适应样本-特征比极低的情况。其线性结构保障了模型输出的可解释性,适用于生物标志物识别。该模型在两个神经退行性疾病数据集上测试,性能超越当前最优模型,并识别出与已有文献报道一致的生物标记特征。
原文摘要 · Abstract (English)
We present BALDUR, a novel Bayesian algorithm designed to deal with multi-modal datasets and small sample sizes in high-dimensional settings while providing explainable solutions. To do so, the proposed model combines within a common latent space the different data views to extract the relevant information to solve the classification task and prune out the irrelevant/redundant features/data views. Furthermore, to provide generalizable solutions in small sample size scenarios, BALDUR efficiently integrates dual kernels over the views with a small sample-to-feature ratio. Finally, its linear nature ensures the explainability of the model outcomes, allowing its use for biomarker identification. This model was tested over two different neurodegeneration datasets, outperforming the state-of-the-art models and detecting features aligned with markers already described in the scientific literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。