arXiv:2606.04437cs.CV2026-06

让自动驾驶车通过精准提问,高效获取异构伙伴的局部感知证据。

INTACT: Ego-Guided Typed Sparse Evidence Retrieval for Heterogeneous Collaborative Perception

论文配图:INTACT: Ego-Guided Typed Sparse Evidence Retrieval for Heterogeneous Collaborative Perception
图 1 · 摘自论文原文
  • ego车发出带类型的任务查询,只索取指定位置的证据,不需整体特征对齐
  • 实测在OPV2V-H上达80.1 AP70,通信量压缩至密集传输的1/16
  • 新设备加入无需训练,只需合并模型检查点,适合快速部署

协同感知通过车辆间信息共享扩展了自动驾驶的感知范围,但异构传感器与感知模型使中间特征融合难以规模化部署。现有方法多采用‘先转换’范式:合作方特征需对齐或投影到本车兼容空间后才能融合。这类兼容性协议虽提升固定系统性能,却将部署绑定于合作方特定适配,导致新异构设备接入成本高昂。为此,我们提出INTACT——一种面向异构协同感知的自我引导型稀疏证据检索框架。不同于整体特征转换,INTACT由本车发出带有对象类型和证据缺失区域的查询,合作方仅返回所查询位置的局部证据,本车通过稀疏查询路由筛选并门控写回。该机制将兼容性要求从全局特征可解释性,转变为在本车查询下的局部、类型化响应可比性,实现了零训练的异构插入协议:本车接口仅需训练一次,新合作者通过检查点合并即可接入。在模拟与真实世界异构协同感知基准上进行的大量实验验证了INTACT的有效性与可部署性。在OPV2V-H上,仅增加0.52M参数与18.0 $\log_2$通信量,即达80.1 AP70,相较密集特征传输约压缩16倍;在DAIR-V2X上,实测复杂场景下仍实现43.8 AP50。

原文摘要 · Abstract (English)

Collaborative perception extends the perceptual range of autonomous vehicles by sharing information across agents, but heterogeneous sensors and perception models make intermediate feature fusion difficult to deploy at scale. Existing heterogeneous collaboration methods typically follow a translation-first paradigm: collaborator features must be aligned, adapted, or projected into an ego-compatible space before fusion. Such feature-compatibility contracts improve fixed-system performance, but they couple deployment to collaborator-specific adaptation and make newly joined heterogeneous agents costly to integrate. To address this gap, we propose INTACT, an ego-guided typed sparse evidence retrieval framework for heterogeneous collaborative perception. Instead of translating an entire collaborator feature map, INTACT lets the ego vehicle issue typed evidence queries that express suspected objects and evidence-deficient regions. Collaborators respond only with local evidence at queried locations, and the ego selects useful responses through sparse per-query routing and injects them through gated residual write-back. This changes the compatibility requirement from global feature-map interpretability to local, typed response comparability under ego-issued queries, enabling a zero-training heterogeneous insertion protocol in which the ego interface is trained once and new collaborators join through checkpoint merging. Extensive experiments on simulated and real-world heterogeneous collaborative perception benchmarks validate the effectiveness and deployability of INTACT. On OPV2V-H, INTACT achieves 80.1 AP70 with only 0.52M additional parameters and 18.0 $\log_2$ communication volume, corresponding to about 16$\times$ compression over dense feature transmission. On DAIR-V2X, INTACT achieves 43.8 AP50 under challenging real-world conditions.

协同感知异构融合稀疏检索自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。