通过多源协同训练提升路侧激光雷达的仿真到真实检测性能。
Solution for UCF UrbanTwin V2X-Real Track: Sim-to-Real Urban LiDAR 3D Object Detection

- 构建多源数据融合训练池,分角色协同优化检测
- 在真实测试集上达3D [email protected] 0.4518,真实性得分0.8871
- 适合关注真实场景三维目标检测的科研与工程人员
解决路边激光雷达仿真到现实的差距需应对场景几何、采样密度、回波模式和行人尺度等多重耦合差异。本文提出一种多源协同训练与类别感知融合框架,将数字孪生扫描、扩散重绘扫描、密度稳定扫描及行人形态对齐样本整合为统一训练池,各具互补角色。在统一的DSVT检测架构下,源专用专家分支保持角色分工并优化同一检测目标。推理时,预设类别感知融合路径分别整合几何稳定与校准感知分支(车辆)、采样互补分支(卡车)以及形态一致证据(行人),并通过无标签点云中心融合精炼几何定位。在UrbanTwin V2X-Real隐藏测试集上,系统综合得分0.7421,3D [email protected]达0.4518,真实性得分0.8871。结果表明,数据源间稳定可解释的协作比无约束模型输出聚合更有效。
原文摘要 · Abstract (English)
Bridging the simulation-to-reality gap in roadside LiDAR requires addressing several coupled discrepancies, including scene geometry, sampling density, return patterns, and pedestrian scale. This report presents a multi-source collaborative training and class-aware fusion framework for Sim2Real 3D detection. The method organizes digital-twin scans, diffusion-redrawn scans, density-stabilized scans, and pedestrian morphology-aligned samples into a unified training pool with complementary roles. Within a common DSVT detection formulation, source-specialized expert branches preserve those roles while optimizing for the same detection objective. At inference, a predefined class-aware fusion pathway integrates geometry-stable and calibration-aware branches for vehicles, sampling-complementary branches for trucks, and morphology-consistent evidence for pedestrians. A label-free point-cloud center blend then refines geometric localization. On the UrbanTwin V2X-Real hidden test set, the unified system achieves a combined score of 0.7421, with 3D [email protected] of 0.4518 and a realism score of 0.8871. The results indicate that a stable, interpretable collaboration among data sources is more valuable than unconstrained aggregation of model outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。