跨模态结构保留学习,让不同数据互补增强表征能力
Multimodal Structure Preservation Learning
- 用一种数据的聚类结构指导另一数据的表征学习
- 在基因组与质谱数据中成功恢复出隐藏聚类结构
- 适合多模态数据融合与流行病学分析场景
在实际应用中构建机器学习模型时,数据的可获取性、采集成本和区分能力是关键考量因素。不同模态的数据常捕捉到现象的不同方面,具有互补性。同时,某些数据源包含对其价值至关重要的结构信息。因此,通过匹配另一数据的结构,可提升某一数据类型的实用性。本文提出多模态结构保留学习(MSPL),一种利用一个数据模态提供的聚类结构来增强另一模态数据表征的新方法。我们在合成时间序列数据中验证了其揭示潜在结构的能力,并在全基因组测序与抗菌药物耐药性数据上,借助质谱数据成功恢复了聚类结构,支持流行病学应用。结果表明,MSPL能够将外部结构注入学习特征中,促进异构数据模态间的有益协同。
原文摘要 · Abstract (English)
When selecting data to build machine learning models in practical applications, factors such as availability, acquisition cost, and discriminatory power are crucial considerations. Different data modalities often capture unique aspects of the underlying phenomenon, making their utilities complementary. On the other hand, some sources of data host structural information that is key to their value. Hence, the utility of one data type can sometimes be enhanced by matching the structure of another. We propose Multimodal Structure Preservation Learning (MSPL) as a novel method of learning data representations that leverages the clustering structure provided by one data modality to enhance the utility of data from another modality. We demonstrate the effectiveness of MSPL in uncovering latent structures in synthetic time series data and recovering clusters from whole genome sequencing and antimicrobial resistance data using mass spectrometry data in support of epidemiology applications. The results show that MSPL can imbue the learned features with external structures and help reap the beneficial synergies occurring across disparate data modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。