arXiv:2607.21137cs.CVcs.LG2026-07

为视障者手机导航设计安全路肩分割模型,兼顾精度与实际部署

Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation

论文配图:Safety-oriented sidewalk and road segmentation for smartphone-based assistive navigation
图 1 · 摘自论文原文
  • 构建胸高视角数据集,用合成图像和SAM2伪标签提升分割效果
  • 模型在真实道路误判率低至7.9%,移动端运行达7.4帧/秒
  • 强调评估需结合准确率、安全误判率和手机性能,适合视障导航应用

独立行走对视障人士至关重要,但手机辅助导航需区分可行走的路肩与邻近危险区域。本研究提出一种面向安全的语义分割框架,构建了包含2,752张图像-掩码对的SENSATION-DS数据集,采用九类导航相关标签体系。将外部城市及路肩数据集统一到该标签空间,并通过分阶段目标域适应方法,使用掩码条件合成图像和Segment Anything Model 2(SAM2)伪标签评估五种分割架构。采用mIoU、道路与路肩特异性指标、以“道路误判为路肩”错误率作为虚假安全代理指标,以及Android Open Neural Network Exchange基准测试进行评估。合成增强普遍提升分割精度,而SAM2伪标签更一致地降低道路误判率。UPerNet-MobileNetV3取得最高离线mIoU(0.715 ± 0.006),DeepLabV3Plus-MobileNetV3则实现最低道路误判率(0.079)和最高安卓端运行效率(512x384分辨率下7.383 FPS)。结果表明,辅助路肩感知应综合评估分割精度、虚假安全行为和手机部署可行性,真实价值仍需视障用户验证。该评估支持选择在感知精度、保守错误行为与实际运行时之间平衡的模型。

原文摘要 · Abstract (English)

Independent sidewalk mobility is essential for blind and visually impaired pedestrians (BVIPs), yet smartphone-based assistive navigation requires perception models that distinguish walkable sidewalks from adjacent unsafe regions. This study presents a safety-oriented semantic segmentation framework for future mobile guidance. We introduce SENSATION-DS, a chest-height pedestrian-view dataset with 2,752 image-mask pairs and nine-class navigation-relevant taxonomy. External urban and sidewalk datasets were harmonized to this label space, and five segmentation architectures were evaluated using staged target-domain adaptation with mask-conditioned synthetic images and Segment Anything Model 2 (SAM2) pseudo-labels. Models were assessed using mean Intersection over Union (mIoU), road- and sidewalk-specific metrics, Road-as-Sidewalk Error Rate as a proxy false-safe measure, and Android Open Neural Network Exchange benchmarking. Synthetic augmentation generally improved segmentation accuracy, whereas SAM2 pseudo-labels more consistently reduced Road-as-Sidewalk errors. UPerNet-MobileNetV3 achieved the highest offline mIoU (0.715 +/- 0.006), while DeepLabV3Plus-MobileNetV3 achieved the lowest Road-as-Sidewalk Error Rate (0.079) and highest Android runtime at 512x384 (7.383 FPS). These results show that assistive sidewalk perception should be evaluated jointly by segmentation accuracy, proxy false-safe behavior, and smartphone deployment feasibility, while real-world benefit requires validation with BVIP users. This evaluation supports selecting models that balance accurate perception, conservative error behavior, and practical runtime.

视觉导航语义分割视障辅助移动端部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。