将视频实时转为适配移动场景的语音提示,帮视障者更安全地感知环境。
AMAVA: Adaptive Motion-Aware Video-to-Audio Framework for Visually-Impaired Assistance

- 根据画面运动强度自动切换语音描述或警报声
- 静态场景用语音描述,动态场景优先提醒危险
- 适合视障人士日常导航,降低听觉负担
为视障和低视力人群设计的导航辅助系统常因持续、无差别的反馈导致认知过载。我们提出AMAVA,一种实时视频转音频框架,将手机摄像头捕捉的视频转化为情境相关的音效或文本转语音描述。该框架采用轻量级AI分类模型识别低运动与高运动场景,进而触发不同音频输出:在静态环境中生成语音场景描述以增强情境感知;在高运动情况下优先输出语音警示和环境音效以保障安全。音频生成基于解码器仅有的视觉-语言模型,结合专家混合与跨模态注意力机制实现视觉理解,并融合神经文本转语音和自然声音合成网络。系统通过提示缓存和类别特定限流策略避免听觉杂乱并降低延迟。我们进行了全面评估,包括真实导航实验,对比仅使用盲杖与结合AMAVA的情况,结果显示用户信心和感知安全性显著提升。
原文摘要 · Abstract (English)
Navigational aids for blind and low vision individuals struggle conveying dynamic real-world environments, leading to cognitive overload from continuous, undifferentiated feedback. We present AMAVA, a novel real-time video-to-audio framework that converts mobile device video into contextually relevant sound effects or text-to-speech descriptions. We propose a motion-aware pipeline using a lightweight AI classification model to distinguish between low and high-movement scenes followed by a real-time text-to-audio synthesis pipeline to enhance environmental perception more efficiently. In static environments, AMAVA generates spoken audio scene descriptions for situational awareness. In high-movement situations, it prioritizes safety by delivering sound cues, such as spoken hazard alerts and environmental sound effects. These audio outputs are produced by a decoder-only transformer-based vision-language model with mixture-of-experts and cross-modal attention for visual understanding, in conjunction with neural text-to-speech and natural sound synthesis networks. The proposed framework uses prompt-based caching and category-specific throttling to avoid auditory clutter and minimize latency. We present a comprehensive evaluation of the system, including a real-time navigation study comparing a white cane alone versus with AMAVA, that shows a significant increase in user confidence and perceived safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。