arXiv:2605.26642cs.CV2026-05

无需适配即可与未知配置的协作智能体高效协同感知。

Adaptation-Free Heterogeneous Collaborative Perception with Unseen Agent Configurations

论文配图:Adaptation-Free Heterogeneous Collaborative Perception with Unseen Agent Configurations
图 1 · 摘自论文原文
  • 将辅助端的检测框信息转换为伪鸟瞰图,生成适配自身特征的融合特征。
  • 在64种未见配置下,相对[email protected]提升35.91%,每帧仅需120字节传输。
  • 适合部署于真实开放场景,如车联网中动态加入的异构车辆感知系统。

协同感知通过共享互补观测提升3D目标检测性能,但现有方法多假设协作端编码器配置固定或已知,限制了实际应用。本文考虑一种开放世界场景:部署后可能出现配置未知的辅助智能体,如不同激光雷达束数或编码器结构。为此,我们提出ALF框架,通过将轻量级框级消息升维为本车兼容的辅助特征,实现零适应协同。ALF将辅助端的框级消息转为伪鸟瞰图,结合本车特征中的目标中心线索与场景上下文,合成适配的潜在特征。在V2X-Real数据集上,64种零样本评估案例中,ALF相对最强基线在[email protected]上提升35.91%,且每帧仅需120字节(10Hz下约9.6 Kbps带宽)。

原文摘要 · Abstract (English)

Collaborative perception improves 3D object detection by enabling agents to share complementary observations, but most existing methods assume fixed or known collaborator encoder configurations, limiting deployment in practice. In this work, we consider an open-world setting in which auxiliary agents with unseen configurations may appear after deployment, such as different LiDAR beam counts or encoder architectures. To address this challenge, we propose ALF, a collaborative perception framework that enables zero-adaptation collaboration with unseen agent configurations by lifting lightweight box-level messages into ego-compatible auxiliary features. ALF converts auxiliary box-level messages into pseudo-BEV maps and synthesizes ego-compatible latent features by combining object-centric cues with scene context from the ego feature. On V2X-Real, under a zero-shot evaluation across 64 case studies, ALF outperforms the strongest prior baseline by 35.91% in relative [email protected] while requiring only 120 bytes per agent per frame (approximately 9.6 Kbps bandwidth at 10 Hz).

协同感知异构系统零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。