针对工业场景多摄像头3D感知难题,提出高效实时解决方案。
A Unified 3D Object Perception Framework for Real-Time Outside-In Multi-Camera Systems
- 基于稀疏4D框架,融合世界坐标先验与遮挡感知重识别模块。
- 在AI City挑战赛上达到45.22的HOTA指标,性能领先。
- 支持单块黑墙级显卡并发64路视频流,适合大规模部署。
精准的3D物体感知与多目标多相机(MTMC)跟踪是工业基础设施数字化的核心。然而,将自动驾驶的“内视”模型迁移到静态摄像头网络的“外视”场景时,面临相机布局异质性和极端遮挡的挑战。本文提出一种专为大型基础设施环境优化的Sparse4D适配框架,利用绝对世界坐标几何先验,并引入遮挡感知的ReID嵌入模块以保持跨分布式传感器网络的身份一致性。为缩小仿真到真实(Sim2Real)域差距且无需人工标注,采用NVIDIA COSMOS框架生成多样环境风格的合成数据,增强模型外观不变性。在AI City Challenge 2025基准上,纯摄像头系统取得45.22的最优HOTA分数。针对实时部署限制,开发了多尺度可变形聚合(MSDA)的TensorRT加速插件,硬件加速后实现2.15倍速度提升,使单块黑墙级GPU可支持超过64路并发摄像头流。
原文摘要 · Abstract (English)
Accurate 3D object perception and multi-target multi-camera (MTMC) tracking are fundamental for the digital transformation of industrial infrastructure. However, transitioning "inside-out" autonomous driving models to "outside-in" static camera networks presents significant challenges due to heterogeneous camera placements and extreme occlusion. In this paper, we present an adapted Sparse4D framework specifically optimized for large-scale infrastructure environments. Our system leverages absolute world-coordinate geometric priors and introduces an occlusion-aware ReID embedding module to maintain identity stability across distributed sensor networks. To bridge the Sim2Real domain gap without manual labeling, we employ a generative data augmentation strategy using the NVIDIA COSMOS framework, creating diverse environmental styles that enhance the model's appearance-invariance. Evaluated on the AI City Challenge 2025 benchmark, our camera-only framework achieves a state-of-the-art HOTA of $45.22$. Furthermore, we address real-time deployment constraints by developing an optimized TensorRT plugin for Multi-Scale Deformable Aggregation (MSDA). Our hardware-accelerated implementation achieves a $2.15\times$ speedup on modern GPU architectures, enabling a single Blackwell-class GPU to support over 64 concurrent camera streams.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。