arXiv:2602.16385cs.CV2026-02被引 1

用单目摄像头提升视障者室内导航,精准识别小障碍物

Adaptive Multi-Scale Channel-Spatial Attention Aggregation Framework for 3D Indoor Semantic Scene Completion Toward Assisting Visually Impaired

  • 设计多尺度注意力融合框架,动态优化特征重建
  • 小物体识别准确率提升16.9%,桌子识别提升10.4%
  • 可部署于嵌入式设备,实现实时辅助导航

视障人士独立室内移动仍面临重大挑战,主要源于现有辅助系统对椅子、桌子等细粒度危险物体检测能力有限,导致陌生环境中碰撞风险显著升高。为弥合单目3D视觉研究与实际辅助应用之间的差距,本文提出一种基于可穿戴RGB相机的自适应多尺度注意力聚合(AMAA)框架,用于单目3D语义场景补全。该框架解决2D到3D特征提升中的两大问题:反投影过程中的噪声扩散和多尺度融合的结构不稳定性。引入并行通道-空间注意力机制,在语义与几何维度上重新校准提升特征;采用分层自适应门控策略,调控跨尺度信息流动,保留细粒度结构细节。在NYUv2基准测试中,AMAA整体mIoU达27.88%。关键的是,相比MonoScene基线,小物体识别相对提升16.9%,桌子识别提升10.4%。此外,基于NVIDIA Jetson Orin NX与ZED~2i相机的可穿戴原型在室内环境中实现稳定实时性能,验证了单目3D场景补全在辅助导航中的可行性。

原文摘要 · Abstract (English)

Independent indoor mobility remains a critical challenge for individuals with visual impairments, largely due to the limited capability of existing assistive systems in detecting fine-grained hazardous objects such as chairs, tables, and small obstacles. These perceptual blind zones substantially increase the risk of collision in unfamiliar environments. To bridge the gap between monocular 3D vision research and practical assistive deployment, this paper proposes an Adaptive Multi-scale Attention Aggregation (AMAA) framework for monocular 3D semantic scene completion using only a wearable RGB camera. The proposed framework addresses two major limitations in 2D-to-3D feature lifting: noise diffusion during back-projection and structural instability in multi-scale fusion. A parallel channel--spatial attention mechanism is introduced to recalibrate lifted features along semantic and geometric dimensions, while a hierarchical adaptive gating strategy regulates cross-scale information flow to preserve fine-grained structural details. Experiments on the NYUv2 benchmark demonstrate that AMAA achieves an overall mIoU of 27.88%. Crucially, it yields significant relative improvements of 16.9% for small objects and 10.4% for tables over the MonoScene baseline. Furthermore, a wearable prototype based on an NVIDIA Jetson Orin NX and a ZED~2i camera validates stable real-time performance in indoor environments, demonstrating the feasibility of deploying monocular 3D scene completion for assistive navigation.

视障辅助3D补全注意力机制可穿戴

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。