用动态代理网络提升多模态推理效率与协作能力
DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning

- 将多模态融合重构为动态代理协作机制,按需激活专用代理
- 在5个数据集上实现新最优,ADNI任务准确率提升2.6%
- 支持上下文感知的可解释协作模式,计算高效且适配复杂场景
当前基于静态专家混合(MoE)架构的多模态融合方法,在复杂现实应用中难以实现自适应、高效的协同推理。本文提出动态代理交互网络(DAIN),将多模态融合重构为动态多代理协作过程。DAIN采用上下文感知的元控制器,动态调度稀疏激活的专用交互代理,并压缩代理间通信以达成共识。框架通过多目标损失函数联合优化任务准确率、代理专业化和运行效率,结合稀疏激活与通信正则化。在五个多样化基准测试——ADNI、MIMIC-IV、MM-IMDB、CMU-MOSI 和 ENRICO 上的全面评估表明,DAIN达到新状态水平,其中在ADNI上取得2.6%的准确率提升。消融实验验证了动态调度和代理通信的关键作用。此外,DAIN通过暴露上下文相关的代理角色与协作模式,提升了可解释性,同时通过样本级稀疏代理激活保持计算效率。本工作展示了动态代理范式在多模态推理中的潜力。
原文摘要 · Abstract (English)
Current multimodal fusion approaches, particularly those based on static Mixture-of-Experts (MoE) architectures, often struggle to provide the adaptive and efficient collaborative reasoning required by complex real-world applications. We introduce the Dynamic Agent-based Interaction Network (DAIN), which reconceptualizes multimodal fusion as a dynamic, multi-agent collaborative process. DAIN employs a context-aware Meta-Controller that dynamically schedules sparse activation of specialized interaction agents and orchestrates compressed inter-agent communication for consensus-building. The framework is guided by a multi-objective loss function that jointly optimizes task accuracy, agent specialization, and operational efficiency through sparse activation and communication regularization. Comprehensive evaluations across five diverse benchmarks -- ADNI, MIMIC-IV, MM-IMDB, CMU-MOSI, and ENRICO -- establish DAIN as a new state-of-the-art, delivering significant performance improvements including a 2.6\% accuracy gain on ADNI. Ablation studies verify the critical roles of both dynamic scheduling and agent communication. Furthermore, DAIN offers enhanced interpretability by exposing context-dependent agent roles and collaboration patterns while maintaining computational efficiency through sample-wise sparse agent activation. Our work demonstrates the promise of dynamic, agent-based paradigms for multimodal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。