从信息论角度分析多模态自编码器的融合机制与效果
Analyzing Multimodal Integration in the Variational Autoencoder from an Information-Theoretic Perspective
- 引入信息论指标,量化各模态对重建的重要性
- 不同权重调度下,模态融合能力差异显著
- 适合研究多模态学习与机器人感知的读者
人类感知本质上是多模态的,例如将视觉、本体感觉和触觉信息整合为统一体验。因此,多模态学习对构建能稳健交互真实世界的机器人系统至关重要。本文研究多模态变分自编码器(Multimodal VAE)在信息融合中的作用。该模型通过编码器将多源输入映射至随机潜在空间,并通过解码器重构数据。在潜在空间中,不同模态在两个时间点进行融合,可作为机器人代理的控制器。本文引入两类信息论度量:单模态误差评估单一模态对自身或全部模态重建的重要性;精度损失则衡量缺失某一模态信息对重建结果的影响。模型基于证据下界(ELBO)训练,其目标函数包含重建项与潜在空间损失项。通过四个不同的潜在损失权重调度策略进行训练,分析其在多模态集成能力上的表现。
原文摘要 · Abstract (English)
Human perception is inherently multimodal. We integrate, for instance, visual, proprioceptive and tactile information into one experience. Hence, multimodal learning is of importance for building robotic systems that aim at robustly interacting with the real world. One potential model that has been proposed for multimodal integration is the multimodal variational autoencoder. A variational autoencoder (VAE) consists of two networks, an encoder that maps the data to a stochastic latent space and a decoder that reconstruct this data from an element of this latent space. The multimodal VAE integrates inputs from different modalities at two points in time in the latent space and can thereby be used as a controller for a robotic agent. Here we use this architecture and introduce information-theoretic measures in order to analyze how important the integration of the different modalities are for the reconstruction of the input data. Therefore we calculate two different types of measures, the first type is called single modality error and assesses how important the information from a single modality is for the reconstruction of this modality or all modalities. Secondly, the measures named loss of precision calculate the impact that missing information from only one modality has on the reconstruction of this modality or the whole vector. The VAE is trained via the evidence lower bound, which can be written as a sum of two different terms, namely the reconstruction and the latent loss. The impact of the latent loss can be weighted via an additional variable, which has been introduced to combat posterior collapse. Here we train networks with four different weighting schedules and analyze them with respect to their capabilities for multimodal integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。