提出新框架提升多模态模型在不确定环境下的鲁棒性
Distributionally Robust Multimodal Machine Learning
- 基于分布鲁棒优化,避免早期融合与盲目不确定性建模
- 理论证明泛化上界与极小极大下界,提供性能保证
- 实测在模拟和真实数据中均显著提升鲁棒性
我们研究分布鲁棒的多模态机器学习问题。现有方法常依赖特征级融合(早期融合)或启发式不确定性建模,忽视模态感知效应且洞察有限。本文提出一种新的分布鲁棒优化(DRO)框架,旨在揭示多模态学习的理论与实践洞见。首先通过复杂性分析验证该设定的重要性;随后建立泛化上界与极小极大下界,提供性能保障。这些结果进一步拓展至考虑编码器特异性误差传播的情形。实验表明,该方法在仿真与真实数据集上均提升了鲁棒性。整体成果为高风险场景中多模态模型的应用提供了原则性基础。
原文摘要 · Abstract (English)
We consider the problem of distributionally robust multimodal machine learning. Existing approaches often rely on merging modalities on the feature level (early fusion) or heuristic uncertainty modeling, which downplays modality-aware effects and provide limited insights. We propose a novel distributionally robust optimization (DRO) framework that aims to study both the theoretical and practical insights of multimodal machine learning. We first justify this setup and show the significance of this problem through complexity analysis. We then establish both generalization upper bounds and minimax lower bounds which provide performance guarantees. These results are further extended in settings where we consider encoder-specific error propogations. Empirically, we demonstrate that our approach improves robustness in both simulation settings and real-world datasets. Together, these findings provide a principled foundation for employing multimodal machine learning models in high-stakes applications where uncertainty is unavoidable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。