通过几何约束提升多模态融合多样性,避免模态主导或信息漂移。
Diverse via bounded Agreement: Geometric Regularization for Multimodal Fusion
- 基于有限共识原则,软约束跨模态过度偏离,保留模态特异性差异。
- 在音视频、图文、射频等任务中,多模态性能普遍提升,单模态表征也增强。
- 轻量级插件式设计,无需修改结构,推理无额外开销,适用广泛。
多模态融合常被视为优化平衡问题,通过调整训练信号防止某一模态主导其他模态。然而,平衡优化无法完全决定中间表示的几何结构。监督式多模态模型仍可能学习低多样性的模态特定嵌入,或导致成对跨模态观测过度分离,削弱单模态鲁棒性与多模态融合效果。本文提出 egName,一种轻量级、可插拔的几何正则化框架,用于多模态表示学习。不同于强制刚性跨模态对齐, egName 遵循有限共识原则:在保留模态特异性多样性的同时,仅软性约束超出可接受一致带宽的成对跨模态漂移。其操作上结合了抑制谱集中度的离散项与控制过度成对漂移的一致带锚定项,无需架构修改或推理时开销。在音频-视觉、图像-文本及基于射频的基准测试中, egName 均显著提升多模态性能,并常增强单模态表示。结果表明,显式调节表示几何是优化平衡的有效补充,且几何感知正则化可在多种架构与领域中改善多模态学习。
原文摘要 · Abstract (English)
Multimodal fusion is often treated as an optimization-balancing problem, where training signals are adjusted to prevent one modality from dominating the others. However, balanced optimization does not fully determine the geometry of intermediate representations. Supervised multimodal models may still learn low-diversity modality-specific embeddings or allow paired cross-modal observations to drift excessively apart, weakening both unimodal robustness and multimodal fusion. We introduce \regName, a lightweight plug-and-play geometric regularization framework for multimodal representation learning. Rather than enforcing rigid cross-modal alignment, \regName follows a bounded-agreement principle: preserve modality-specific diversity while softly constraining only the portion of paired cross-modal drift that exceeds an admissible agreement band. Operationally, \regName combines a dispersion term that mitigates spectral concentration with an agreement-band anchoring term that controls excessive paired drift, requiring no architectural modification or inference-time overhead. Experiments across audio-visual, image-text, and RF-based benchmarks show that \regName consistently improves multimodal performance and often strengthens unimodal representations. These results suggest that explicitly regulating representation geometry is an effective complement to optimization balancing, and provide evidence that geometry-aware regularization can improve multimodal learning across diverse architectures and domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。