多智能体协作补全缺失模态,提升自动驾驶与机器人决策能力
CAML: Collaborative Auxiliary Modality Learning for Multi-Agent Systems
- 多智能体共享训练数据,推理时可缺损模态仍能协同工作
- 自动驾驶场景下事故检测准确率提升58.1%,语义分割mIoU提升10.6%
- 适用于车路协同、无人机地面机器人等动态多模态环境
多模态学习在自动驾驶、机器人和推理等领域已成为提升性能的关键技术。然而,在资源受限环境下,训练时可用的某些模态在推理阶段可能缺失。现有方法虽能在训练中利用多源数据并支持推理时减少模态,但主要针对单智能体场景。这对车联网等动态环境中的连通自动驾驶车辆(CAV)构成挑战,因数据覆盖不全易导致决策盲区。相反,部分多智能体研究未解决测试时模态缺失问题。为此,本文提出协同辅助模态学习(CAML),一种新型多模态多智能体框架:训练时智能体协作共享多模态数据,推理时允许模态缺失。在事故高发场景下的协同决策实验中,CAML实现事故检测最高58.1%的性能提升;在真实世界空地机器人数据上进行协同语义分割,mIoU最高提升10.6%。
原文摘要 · Abstract (English)
Multi-modal learning has emerged as a key technique for improving performance across domains such as autonomous driving, robotics, and reasoning. However, in certain scenarios, particularly in resource-constrained environments, some modalities available during training may be absent during inference. While existing frameworks effectively utilize multiple data sources during training and enable inference with reduced modalities, they are primarily designed for single-agent settings. This poses a critical limitation in dynamic environments such as connected autonomous vehicles (CAV), where incomplete data coverage can lead to decision-making blind spots. Conversely, some works explore multi-agent collaboration but without addressing missing modality at test time. To overcome these limitations, we propose Collaborative Auxiliary Modality Learning (CAML), a novel multi-modal multi-agent framework that enables agents to collaborate and share multi-modal data during training, while allowing inference with reduced modalities during testing. Experimental results in collaborative decision-making for CAV in accident-prone scenarios demonstrate that CAML achieves up to a 58.1% improvement in accident detection. Additionally, we validate CAML on real-world aerial-ground robot data for collaborative semantic segmentation, achieving up to a 10.6% improvement in mIoU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。