arXiv:2410.06395cs.LGcs.AI2024-10被引 4

自动构建图结构,让模型能灵活处理任意数量的多模态数据。

Multimodal Representation Learning using Adaptive Graph Construction

  • 通过图优化自动构造多模态连接,无需手动设计架构。
  • 在阿尔茨海默病检测任务中超越现有方法,准确率提升显著。
  • 适合需要融合多种数据源的医疗、跨模态分析场景。

多模态对比学习通过利用图像、文本等异构数据训练神经网络。然而,许多现有架构无法泛化到任意数量的模态,需人工设计。本文提出AutoBIND,一种基于图优化的新型对比学习框架,可自动从任意数量的模态中学习表示。我们在阿尔茨海默病检测任务上评估该方法,因其具有实际医学应用价值且涵盖多种数据模态。结果表明,AutoBIND在该任务上优于先前方法,验证了其通用性。

原文摘要 · Abstract (English)

Multimodal contrastive learning train neural networks by levergaing data from heterogeneous sources such as images and text. Yet, many current multimodal learning architectures cannot generalize to an arbitrary number of modalities and need to be hand-constructed. We propose AutoBIND, a novel contrastive learning framework that can learn representations from an arbitrary number of modalites through graph optimization. We evaluate AutoBIND on Alzhiemer's disease detection because it has real-world medical applicability and it contains a broad range of data modalities. We show that AutoBIND outperforms previous methods on this task, highlighting the generalizablility of the approach.

多模态学习图神经网络对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。