arXiv:2606.07708cs.CVcs.AI2026-06被引 1

用无人机视频监督,实现单目视角到鸟瞰图的精准定位与跨视角目标匹配。

Cross-View Urban Traffic Dataset: Drone-Supervised Ground Truth for Monocular Bird's-Eye View Localization

论文配图:Cross-View Urban Traffic Dataset: Drone-Supervised Ground Truth for Monocular Bird's-Eye View Localization
图 1 · 摘自论文原文
  • 通过同步自行车与无人机视频,建立跨视角目标追踪与坐标转换的新基准。
  • 跨视角匹配召回率高但存在误分配和时间不一致问题,鸟瞰图预测仍有提升空间。
  • 适合研究城市交通感知、多视角对齐及自动驾驶全局理解的学者使用。

我们提出一个用于跨视角城市交通感知的数据集与评测基准,基于真实城市交叉口同步采集的自车视角自行车视频与航拍无人机视频。该基准聚焦两个关联任务:街景与航拍视角间的目标轨迹身份匹配,以及利用航拍监督实现自车到鸟瞰图(BEV)的预测。相比以往的城市驾驶与车路协同数据集,本基准提供跨视角的身份级对齐、标准化评估、标注工具及基线实现。该设定源于以交叉口为中心的交通分析需求,需在不同视角下联合推理目标身份、局部交互与全局空间结构。我们在轨迹与帧级别评估方法,包括跨视角ID精确率/召回率/IDF1、远近区分、时间稳定性与一致性指标。同时提供楔形匹配基线及三种BEV预测基线:逆透视映射、类MonoLayout的可学习基线与回归基线。结果表明该基准可行但具挑战性:跨视角匹配虽有高召回率,但受过分配与时间不一致限制;自车到鸟瞰图预测得益于航拍监督,但在轻量单目感知下仍未饱和。期望该基准推动未来跨视角感知、场景对齐与全局交通理解的研究。

原文摘要 · Abstract (English)

We introduce a dataset and benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections. The benchmark targets two linked tasks: cross-view identity matching between street-view and drone-view object tracks, and ego-to-bird's-eye-view prediction using aerial supervision. In contrast to prior urban driving and V2X datasets, our benchmark provides identity-level alignment across radically different viewpoints together with standardized evaluation, annotation tooling, and baseline implementations. This setting is motivated by intersection-centric traffic analysis, where identity preservation, local interactions, and global spatial structure must be reasoned about jointly across views. We evaluate methods at both the track and frame levels, including cross-view ID precision/recall/IDF1, near--far breakdowns, temporal stability, and consistency metrics. We also provide baseline results for wedge-based cross-view matching and for three BEV prediction baselines: inverse perspective mapping, a MonoLayout-style learned baseline, and a regression baseline. The results show that the benchmark is feasible but challenging: cross-view matching achieves strong recall yet remains limited by over-assignment and temporal inconsistency, while ego-to-BEV prediction benefits from aerial supervision but remains far from saturated under lightweight monocular sensing. We hope that this benchmark will support future research on cross-view perception, urban scene alignment, and ego-to-global traffic understanding.

跨视角感知鸟瞰图预测城市交通数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。