arXiv:2511.02946cs.CV2025-11被引 1

ProM3E可任意生成生态多模态嵌入,支持模态缺失重建与融合决策。

ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology

  • 基于嵌入空间的掩码模态重建,实现跨模态信息补全。
  • 在跨模态检索中综合内外部相似性,性能全面领先。
  • 具备模态融合可行性分析能力,适合生态多源数据研究者。

我们提出ProM3E,一种用于生态学任意模态间生成的概率化掩码多模态嵌入模型。该模型基于嵌入空间中的掩码模态重建,能够根据部分上下文模态推断缺失模态。设计上支持嵌入空间中的模态反转。其概率特性使我们能分析不同模态组合在特定下游任务中的融合可行性,本质上学习应融合哪些模态。利用这些特性,我们提出一种新颖的跨模态检索方法,融合模态间与模态内相似性,在所有检索任务中均取得优异表现。此外,我们利用模型隐含表示进行线性探测,验证了其强大的表征学习能力。所有代码、数据集及模型将公开于 https://vishu26.github.io/prom3e。

原文摘要 · Abstract (English)

We introduce ProM3E, a probabilistic masked multimodal embedding model for any-to-any generation of multimodal representations for ecology. ProM3E is based on masked modality reconstruction in the embedding space, learning to infer missing modalities given a few context modalities. By design, our model supports modality inversion in the embedding space. The probabilistic nature of our model allows us to analyse the feasibility of fusing various modalities for given downstream tasks, essentially learning what to fuse. Using these features of our model, we propose a novel cross-modal retrieval approach that mixes inter-modal and intra-modal similarities to achieve superior performance across all retrieval tasks. We further leverage the hidden representation from our model to perform linear probing tasks and demonstrate the superior representation learning capability of our model. All our code, datasets and model will be released at https://vishu26.github.io/prom3e.

多模态生态学嵌入模型生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。