arXiv:2602.08597cs.AI2026-02中稿 · ICANN 2026

用轻量级注意力机制提升多模态系统在噪声下的鲁棒性

An Attention Mechanism for Robust Multimodal Integration in a Global Workspace Architecture

  • 基于全局工作区理论设计顶层注意力选择器,独立于特征提取
  • 在两种数据集上显著提升抗干扰能力,参数量仅为端到端方法的1/10
  • 策略可跨任务、跨扰动类型迁移,甚至适配新模态

稳健的多模态系统需在部分模态出现噪声、退化或不可靠时仍保持性能。现有融合方法常将模态选择与表征学习联合训练,难以判断鲁棒性源于选择器本身还是端到端协同适应。受全局工作区理论(GWT)启发,本文采用轻量级自上而下的模态选择器,作用于冻结的多模态全局工作区之上。我们在两个复杂度递增的多模态数据集Simple Shapes和MM-IMDb 1.0上评估该方法,施加结构化模态扰动。结果表明,该选择器在使用远少于端到端注意力基线的可训练参数下,显著提升系统鲁棒性;所学选择策略在下游任务、扰动模式间具有更强泛化能力,甚至可应用于未见过的模态。此外,在MM-IMDb 1.0基准上,该机制在无显式扰动条件下亦优于无注意力的全局工作区,达到良好基准表现。

原文摘要 · Abstract (English)

Robust multimodal systems must remain effective when some modalities are noisy, degraded, or unreliable. Existing multimodal fusion methods often learn modality selection jointly with representation learning, making it difficult to determine whether robustness comes from the selector itself or from full end-to-end co-adaptation. Motivated by Global Workspace Theory (GWT), we study this question using a lightweight top-down modality selector operating on top of a frozen multimodal global workspace. We evaluate our method on two multimodal datasets of increasing complexity: Simple Shapes and MM-IMDb 1.0, under structured modality corruptions. The selector improves robustness while using far fewer trainable parameters than end-to-end attention baselines, and the learned selection strategy transfers better across downstream tasks, corruption regimes, and even to a previously unseen modality. Beyond explicit corruption settings, on the MM-IMDb 1.0 benchmark, we show that the same mechanism improves the global workspace over its no-attention counterpart and yields decent benchmark performance.

多模态注意力机制鲁棒性全局工作区

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。