arXiv:2605.01967cs.LGcs.CV2026-05

提出新方法提升多模态模型跨环境泛化能力

MER-DG: Modality-Entropy Regularization for Multimodal Domain Generalization

论文配图:MER-DG: Modality-Entropy Regularization for Multimodal Domain Generalization
图 1 · 摘自论文原文
  • 通过最大化各模态特征熵,防止模型依赖特定场景的模态共现关系
  • 在EPIC-Kitchens和HAC数据集上比标准融合方法提升约5%
  • 可无缝接入现有框架,适合需跨场景部署的多模态系统

将多模态模型应用于真实场景时,需在记录条件不同的新环境中实现泛化,这被称为多模态领域泛化(MMDG)。现有方法通常为每种模态使用独立编码器,并通过融合模块端到端训练。本文发现,这种联合优化会导致编码器利用模态间的共现统计关系(由源域特定录制条件引发),而非学习域不变特征,称之为融合过拟合。为此,我们提出用于领域泛化的模态熵正则化(MER-DG),通过最大化每个编码器特征分布的熵来保持特征多样性。MER-DG具有架构无关性,可作为附加损失项融入现有多模态框架。在EPIC-Kitchens和HAC基准上的大量实验表明,其平均性能比标准融合方法提升约5%,比当前最先进方法提升约2%。

原文摘要 · Abstract (English)

Deploying multimodal models in real-world scenarios requires generalization to new environments where recording conditions differ from training, a challenge known as multimodal domain generalization (MMDG). Standard architectures employ separate encoders for each modality and a fusion module, training the system end-to-end by optimizing on the fused features. In this paper, we identify that such joint optimization causes encoders to exploit cross-modal co-occurrences, statistical relationships between modalities that arise from source-specific recording conditions, rather than learning domain-invariant features. We term this failure mode Fusion Overfitting. To address this, we propose Modality-Entropy Regularization for Domain Generalization (MER-DG), which maximizes the entropy of each encoder's feature distribution to preserve feature diversity. MER-DG is architecture-agnostic and integrates into existing multimodal frameworks as an additive loss term. Extensive experiments on EPIC-Kitchens and HAC benchmarks demonstrate average improvements of approximately 5% over standard fusion and approximately 2% over state-of-the-art methods.

多模态域泛化熵正则特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。