arXiv:2503.19564cs.LGcs.AI2025-03被引 5

联邦多模态学习框架,提升动态环境下的可信与可解释性。

FedMM-X: A Trustworthy and Interpretable Framework for Federated Multi-Modal Learning in Dynamic Environments

  • 融合联邦学习与可解释多模态推理,统一跨模态一致性验证。
  • 在视觉语言任务中提升准确率,降低对抗攻击与虚假关联风险。
  • 支持动态客户端参与,量化全局模型可靠性,适合高可信场景。

随着人工智能系统在真实环境中广泛应用,视觉、语言、音频等多模态数据的融合带来了前所未有的机遇与挑战,尤其在实现可信智能方面。本文提出一种新框架 FedMM-X(联邦多模态可解释智能),将联邦学习与可解释多模态推理相结合,以应对去中心化动态环境中的数据异构性、模态不平衡和分布外泛化问题。该框架引入跨模态一致性检测、客户端级可解释机制以及动态信任校准策略。在包含视觉-语言任务的联邦多模态基准上进行严格评估,结果表明其在准确率和可解释性方面均有提升,同时降低了对对抗样本和虚假相关性的脆弱性。此外,提出一种新颖的信任得分聚合方法,用于量化动态客户端参与下全局模型的可靠性。研究为构建鲁棒、可解释且符合社会伦理的现实世界AI系统提供了路径。

原文摘要 · Abstract (English)

As artificial intelligence systems increasingly operate in Real-world environments, the integration of multi-modal data sources such as vision, language, and audio presents both unprecedented opportunities and critical challenges for achieving trustworthy intelligence. In this paper, we propose a novel framework that unifies federated learning with explainable multi-modal reasoning to ensure trustworthiness in decentralized, dynamic settings. Our approach, called FedMM-X (Federated Multi-Modal Explainable Intelligence), leverages cross-modal consistency checks, client-level interpretability mechanisms, and dynamic trust calibration to address challenges posed by data heterogeneity, modality imbalance, and out-of-distribution generalization. Through rigorous evaluation across federated multi-modal benchmarks involving vision-language tasks, we demonstrate improved performance in both accuracy and interpretability while reducing vulnerabilities to adversarial and spurious correlations. Further, we introduce a novel trust score aggregation method to quantify global model reliability under dynamic client participation. Our findings pave the way toward developing robust, interpretable, and socially responsible AI systems in Real-world environments.

联邦学习多模态可解释性可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。