arXiv:2503.16488cs.HCcs.CV2025-03被引 9

用视觉语言模型和距离感知检测,实时为视障者生成带距离信息的语音描述。

VocalEyes: Enhancing Environmental Perception for the Visually Impaired through Vision-Language Models and Distance-Aware Object Detection

  • 用4比特量化微调Florence-2模型,在边缘设备上实现低延迟视频理解。
  • 5帧延迟内输出物体、行人、障碍物及其距离的语音描述,准确率高。
  • 支持34种音色自定义,适合需要个性化语音反馈的视障用户。

随着对提升视障人士独立性与移动能力的辅助技术需求增加,本文提出一种创新的实时系统,通过音频描述增强用户环境感知能力。系统接收实时视频输入,利用经过4比特量化与微调的Florence-2大模型,在NVIDIA Jetson Orin Nano等低功耗边缘设备上高效运行。模型以5帧延迟将视频信号转换为包含物体、行人、障碍物及其估计距离的上下文相关描述。系统采用轻量级可定制的Parler TTS Mini进行语音反馈,支持34种不同说话人类型,并可调节语调、语速与风格以适配用户需求。研究探讨了用于该应用的量化与微调技术,证明结合紧凑模型架构与多功能TTS组件可显著提升实时性能与用户体验。系统在准确性、效率与实用性方面均表现良好,为视障者安全导航提供可行方案。

原文摘要 · Abstract (English)

With an increasing demand for assistive technologies that promote the independence and mobility of visually impaired people, this study suggests an innovative real-time system that gives audio descriptions of a user's surroundings to improve situational awareness. The system acquires live video input and processes it with a quantized and fine-tuned Florence-2 big model, adjusted to 4-bit accuracy for efficient operation on low-power edge devices such as the NVIDIA Jetson Orin Nano. By transforming the video signal into frames with a 5-frame latency, the model provides rapid and contextually pertinent descriptions of objects, pedestrians, and barriers, together with their estimated distances. The system employs Parler TTS Mini, a lightweight and adaptable Text-to-Speech (TTS) solution, for efficient audio feedback. It accommodates 34 distinct speaker types and enables customization of speech tone, pace, and style to suit user requirements. This study examines the quantization and fine-tuning techniques utilized to modify the Florence-2 model for this application, illustrating how the integration of a compact model architecture with a versatile TTS component improves real-time performance and user experience. The proposed system is assessed based on its accuracy, efficiency, and usefulness, providing a viable option to aid vision-impaired users in navigating their surroundings securely and successfully.

视障辅助多模态感知边缘计算语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。