arXiv:2605.00370cs.LGcs.CY2026-05中稿 · ICML被引 3

通过两阶段协作机制,解决多模态模型的主导与虚假关联问题。

Group Cognition Learning: Making Everything Better Through Governed Two-Stage Agents Collaboration

论文配图:Group Cognition Learning: Making Everything Better Through Governed Two-Stage Agents Collaboration
图 1 · 摘自论文原文
  • 设计路由与审计双代理,动态筛选有效跨模态交互。
  • 在三个数据集上实现回归与分类任务的最新性能。
  • 适合多模态融合、情感分析等需要平衡模态贡献的研究者。

集中式多模态学习通常将语言、声学和视觉信号压缩为单一融合表示进行预测。尽管有效,但该范式存在两个局限:模态主导性,即优化偏向阻力最小路径,忽略较弱但有信息量的模态;以及虚假模态耦合,即模型过拟合于偶然的跨模态相关性。为解决此问题,我们提出群体认知学习(GCL),一种受控协作范式,在各模态独立编码后采用两阶段协议。第一阶段(选择性交互)中,路由代理提出定向交互路径,审计代理为每个样本分配门控,强调带来正边际预测增益的交换,抑制冗余耦合。第二阶段(共识形成)中,公共因子代理维护显式共享因子,聚合代理通过贡献感知加权生成最终预测,同时保留各模态表示作为专业化通道。在CMU-MOSI、CMU-MOSEI和MIntRec上的大量实验表明,GCL有效缓解了模态主导与虚假耦合,在回归与分类基准上均达到最先进水平。分析实验进一步验证了设计的有效性。

原文摘要 · Abstract (English)

Centralized multimodal learning commonly compresses language, acoustic, and visual signals into a single fused representation for prediction. While effective, this paradigm suffers from two limitations: modality dominance, where optimization gravitates towards the path of least resistance, ignoring weaker but informative modalities, and spurious modality coupling, where models overfit to incidental cross-modal correlations. To address these, we propose Group Cognition Learning (GCL), a governed collaboration paradigm that applies a two-stage protocol after modality-specific encoding. In Stage 1 (Selective Interaction), a Routing Agent proposes directed interaction routes, and an Auditing Agent assigns sample-wise gates to emphasize exchanges that yield positive marginal predictive gain while suppressing redundant coupling. In Stage 2 (Consensus Formation), a Public-Factor Agent maintains an explicit shared factor, and an Aggregation Agent produces the final prediction through contribution-aware weighting while keeping each modality representation as a specialization channel. Extensive experiments on CMU-MOSI, CMU-MOSEI, and MIntRec demonstrate that GCL mitigates dominance and coupling, establishing state-of-the-art results across both regression and classification benchmarks. Analysis experiments further demonstrate the effectiveness of the design.

多模态学习模态对齐协作机制情感分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。