用几何约束自动识别自行车后视视频中的超车行为,准确率97.8%。
A Geometry-Informed Computer Vision Method for Detecting and Examining Overtaking Vehicles From A Bicycle

- 结合目标检测与轨迹跟踪,通过视角投影原理验证超车三要素
- 97.8%召回率,零误报,可提前2.44秒预判超车意图
- 无需标定即可估算横向距离,误差仅13-14厘米,适合安全预警
车载仪器研究提供了车辆超车行为的直接现场证据,但从连续后视视频中提取超车事件仍依赖人工逐帧标注,成为样本量受限的主要瓶颈。本文提出一种几何感知的计算机视觉流程,仅用单个自行车摄像机即可自动化检测超车事件,无需多传感器配置或相机标定。系统融合RT-DETR目标检测与ByteTrack多目标追踪,通过三阶段几何验证模块,依据透视投影原理强制执行方位角趋势、视觉尺寸增长和空间一致性标准。在密歇根州安娜堡城市道路采集的315个手动标注真实超车事件上测试,系统实现97.8%召回率且无误报。系统平均在车辆通过前2.44秒识别超车意图,84.1%的事件超过1.5秒的人类反应时间阈值,证明其具备主动骑行预警可行性。对96个事件的横向距离测量显示,33.3%的超车距离低于5英尺(152.4厘米)阈值,与以往实地及自报研究一致。初步提出的免标定横向距离估计算法,利用边界框几何特征,在留一交叉验证下平均绝对误差为13–14厘米,足以区分近距离与常规超车,支持安全分类。该系统通过自动化处理消费级影像中的事件提取,消除了仪器化自行车研究的主要标注瓶颈,为更大规模数据集和多样化城市环境下的车-自行车交互分析提供可扩展基础。
原文摘要 · Abstract (English)
Instrumented bicycle studies have produced direct field evidence on vehicle passing behavior, but extracting overtaking events from continuous rear-facing video has remained dependent on manual, frame-by-frame annotation. This bottleneck constrains sample sizes and limits naturalistic cycling safety research. We present a geometry-informed computer vision pipeline that automates overtaking event detection from a single bicycle-mounted camera without multi-sensor configurations or explicit camera calibration. The system combines RT-DETR object detection with ByteTrack multi-object tracking through a three-stage geometric validation module enforcing bearing angle trend, apparent size growth, and spatial confirmation criteria derived from perspective projection principles. Validated on 315 manually annotated real-world overtaking events from urban roads in Ann Arbor, Michigan, the pipeline achieved 97.8% recall with zero false positives. The system identified overtaking intentions a mean of 2.44 seconds before vehicle passage, with 84.1% of events exceeding the 1.5-second human reaction time threshold, demonstrating feasibility for active cyclist warning. Lateral passing distance measurements from 96 events revealed 33.3% of passes below the 5-foot (152.4 cm) threshold, consistent with non-compliance rates in prior field and self-reported studies. A preliminary calibration-free lateral distance estimation approach using bounding box geometric features achieved mean absolute errors of 13-14 cm under leave-one-out cross-validation, sufficient to distinguish close passes from standard passes for safety categorization. By automating event isolation from consumer-grade footage, the system removes the primary annotation bottleneck of instrumented bicycle research and provides a scalable foundation for vehicle-bicycle interaction analysis across larger datasets and diverse urban environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。