arXiv:2609.03818cs.AI2026-09

解决多智能体感知中异构传感器导致的语义不一致问题

CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception

论文配图:CauseCollab: Causal Unified and Modality-Agnostic Network for Heterogeneous Collaborative Perception
图 1 · 摘自论文原文
  • 从因果视角分离语义与模态干扰,提升协议空间表征质量
  • 在OPV2V和DAIR-V2X上达到最优性能,大模态差异场景增益显著
  • 新增模态仅需训练少量适配器,支持灵活扩展

协同感知通过多智能体信息共享增强环境理解,但实际应用中受限于异构传感器模态和模型架构。现有基于协议的两阶段方法虽将异构特征映射至共享协议空间,但独立训练的模态专用转换器常生成模态特异性伪协议分布,导致语义不一致与误差累积,尤其在模态差异大的场景下更明显。为此,我们提出CauseCollab——一种因果统一且模态无关的网络。CauseCollab从因果角度建模协议空间中的表示学习,通过因果度量学习显式解耦语义因素与模态特异性统计混杂因子;同时采用上下文引导的统一转换器,确保跨模态语义一致性。此外,集成新模态仅需训练参数极少的适配器。在OPV2V和DAIR-V2X数据集上的大量实验表明,CauseCollab取得当前最优性能,尤其在存在显著模态差距的场景中表现更优。

原文摘要 · Abstract (English)

Collaborative perception enhances environment understanding through multi-agent information sharing, but its performance in real-world scenarios is constrained by heterogeneous sensor modalities and model architectures. Recent protocol-based two-stage methods alleviate this problem by mapping heterogeneous features into a shared protocol space; however, independently trained modality-specific converters often generate modality-specific pseudo-protocol distributions, leading to semantic inconsistency and error accumulation, which is particularly pronounced in scenarios with large modality discrepancies. To address this issue, we propose CauseCollab, a causal unified and modality-agnostic network. CauseCollab formulates representation learning in the protocol space from a causal perspective, explicitly disentangling semantic factors from modality-specific statistical confounders via causal metric learning. Meanwhile, CauseCollab adopts context-guided Unified Converter for heterogeneous modalities to ensure cross-modal semantic consistency. In addition, integrating new modalities only requires training adapters with minimal parameters. Extensive experiments on the OPV2V and DAIR-V2X datasets demonstrate that CauseCollab achieves state-of-the-art performance, with more significant gains in scenarios involving large modality gaps.

协同感知因果学习异构融合多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。