提出循环自适应修正流,解决多模态融合中的模态差距问题。
CaReFlow: Cyclic Adaptive Rectified Flow for Multimodal Fusion
- 采用一对多映射策略,让源模态数据感知目标模态全局分布。
- 在跨样本对上放松对齐,提升分布映射精度与鲁棒性。
- 引入循环修正流,确保特征可逆迁移,保留模态特异性信息。
模态差距严重限制了多模态融合的效果。现有方法多使用扩散模型或对抗学习进行一一对应对齐,但未让源模态数据充分接触目标模态的全局分布信息。为此,本文基于修正流可沿直线轨迹映射分布的特性,扩展其用于模态分布对齐。具体地,利用修正流中的一对多映射机制,使源模态每个数据点能观测目标模态整体分布,缓解单样本内配对数据不足的问题,实现更稳健的分布变换。此外,设计自适应松弛对齐策略:对同一样本内的模态对施加严格对齐,对跨样本或跨类别对则采用松弛映射,以应对方向模糊性。为防止映射过程中的信息丢失,引入循环修正流,确保转换后的特征可还原为原始特征,从而让多模态表示学习到充分的模态特异性信息。在多种多模态情感计算任务上,仅用简单融合方法即取得优异性能,可视化结果验证了该方法有效缩小模态差距。
原文摘要 · Abstract (English)
Modality gap significantly restricts the effectiveness of multimodal fusion. Previous methods often use techniques such as diffusion models and adversarial learning to reduce the modality gap, but they typically focus on one-to-one alignment without exposing the data points of the source modality to the global distribution information of the target modality. To this end, leveraging the characteristic of rectified flow that can map one distribution to another via a straight trajectory, we extend rectified flow for modality distribution mapping. Specifically, we leverage the `one-to-many mapping' strategy in rectified flow that allows each data point of the source modality to observe the overall target distribution. This also alleviates the issue of insufficient paired data within each sample, enabling a more robust distribution transformation. Moreover, to achieve more accurate distribution mapping and address the ambiguous flow directions in one-to-many mapping, we design `adaptive relaxed alignment', enforcing stricter alignment for modality pairs belonging to the same sample, while applying relaxed mapping for pairs not belonging to the same sample or category. Additionally, to prevent information loss during distribution mapping, we introduce `cyclic rectified flow' to ensure the transferred features can be translated back to the original features, allowing multimodal representations to learn sufficient modality-specific information. After distribution alignment, our approach achieves very competitive results on multiple tasks of multimodal affective computing even with a simple fusion method, and visualizations verify that it can effectively reduce the modality gap.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。