无需标定相机参数,用深度学习从街景图生成道路使用者的鸟瞰矢量图
TopView: Vectorising road users in a bird's eye view from uncalibrated street-level imagery with deep learning
- 通过学习场景消失点实现多视角物体到鸟瞰图的正交投影
- 利用消失点与轨迹线将2D边界框转为3D空间信息,定位精度高
- 适用于城市级社交距离监控与实时地图构建,适合智能交通研究者
生成道路使用者的鸟瞰视图对导航、检测行为冲突、测量空间占用等应用具有重要意义,且可使用度量系统计算物体间距离。本文提出一种无需预先知晓相机内参和外参的简单方法,通过学习场景消失点,将不同视角的物体投影至鸟瞰图。同时,结合学习到的消失点与轨迹线,将道路使用者的2D边界框转换为3D空间信息。该框架已应用于多个场景,包括从摄像头流生成实时地图,以及在城市尺度分析社交距离违规行为。实验表明,该方法在多种未标定摄像头下均能实现道路使用者的高精度地理定位,为城市建模技术的新发展提供了可能,并支持基于深度学习与计算机视觉的建筑环境精准模拟,有助于改进基于代理的建模。
原文摘要 · Abstract (English)
Generating a bird's eye view of road users is beneficial for a variety of applications, including navigation, detecting agent conflicts, and measuring space occupancy, as well as the ability to utilise the metric system to measure distances between different objects. In this research, we introduce a simple approach for estimating a bird's eye view from images without prior knowledge of a given camera's intrinsic and extrinsic parameters. The model is based on the orthogonal projection of objects from various fields of view to a bird's eye view by learning the vanishing point of a given scene. Additionally, we utilised the learned vanishing point alongside the trajectory line to transform the 2D bounding boxes of road users into 3D bounding information. The introduced framework has been applied to several applications to generate a live Map from camera feeds and to analyse social distancing violations at the city scale. The introduced framework shows a high validation in geolocating road users in various uncalibrated cameras. It also paves the way for new adaptations in urban modelling techniques and simulating the built environment accurately, which could benefit Agent-Based Modelling by relying on deep learning and computer vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。