arXiv:2606.00891cs.CV2026-06

首个多模态领域泛化基准,揭示模型组合与模态稳定性的关键关系

MMDG-Bench: A Benchmark for Multimodal Domain Generalization

论文配图:MMDG-Bench: A Benchmark for Multimodal Domain Generalization
图 1 · 摘自论文原文
  • 构建D2M与M2D两种统一框架,系统评估多模态领域泛化效果
  • 在动作识别与人脸反欺诈任务中,新方法显著超越现有最佳性能
  • 发现模态稳定性决定框架选择,强模型在结构化组合中收益更大

多模态领域泛化(MMDG)旨在利用互补模态提升模型在未见域上的鲁棒性。尽管多模态学习(MML)和领域泛化(DG)各自进展迅速,但两者的系统融合仍不充分。当前研究多集中于动作识别,缺乏标准化评估协议。为此,我们提出MMDG-Bench,包含两个基础框架:先领域泛化后多模态学习(D2M)和先多模态学习后领域泛化(M2D)。提供跨视频-音频-光流动作识别与RGB-深度-红外人脸反欺诈等多样任务的统一实验协议。通过在两种框架下组合统一的MML配置与五种DG技术,构建十种MMDG基线,结果表明结构化组合普遍优于现有最优方法。分析揭示三大关键洞察:(1) 集成DG技术可稳定提升泛化性能,非DG方法对骨干网络变化敏感;(2) 框架选择取决于模态间稳定性:模态关系稳定时选D2M,跨域关系差异大时选M2D更鲁棒;(3) 强骨干网络在结构化框架中获得更显著性能增益。MMDG-Bench为未来多模态鲁棒性研究提供原则性基础与实用设计指南。代码已开源:https://github.com/qszhan/MMDG-Bench。

原文摘要 · Abstract (English)

Multi-modal Domain Generalization (MMDG) seeks to leverage complementary modalities to enhance model robustness on unseen domains. Despite extensive progress in Multi-modal Learning (MML) and Domain Generalization (DG) as individual fields, their systematic integration remains under-explored. Current MMDG research is largely confined to action recognition and lacks standardized evaluation protocols. To address this, we introduce MMDG-Bench, a comprehensive benchmark featuring two foundational frameworks: DG then MML (D2M) and MML then DG (M2D). We provide unified experimental protocols across diverse tasks, including video-audio-flow action recognition and RGB-Depth-IR face anti-spoofing. By instantiating ten MMDG baselines through pairing a unified MML configuration with five DG techniques under both D2M and M2D orderings, we demonstrate that these structured combinations frequently outperform existing state-of-the-art methods, underscoring the necessity of a unified benchmarking effort. Our analysis yields three key insights: (1) Integrating DG techniques provides consistent generalization gains across various backbones, whereas non-DG methods are highly sensitive to backbone shifts; (2) The optimal framework choice depends on inter-modal stability: D2M excels when modal relations are stable across domains, while M2D is more robust to cross-domain relational variance; (3) Stronger backbones yield amplified performance dividends when integrated into our structured frameworks. MMDG-Bench provides a principled foundation and actionable design guidelines for future research in multi-modal robustness. Code is released at https://github.com/qszhan/MMDG-Bench.

多模态领域泛化基准测试模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。