Sanvaad让聋人与视障者用手语和语音实现双向实时交流
Sanvaad: A Multimodal Accessibility Framework for ISL Recognition and Voice-Based Interaction
- 基于MediaPipe手部关键点识别手语,轻量高效可在手机运行
- 语音转手语生成动画或字母可视化,支持双向互动
- 为视障者提供无屏语音界面,适合多语言环境使用
聋人与视障者及普通听觉人群之间的交流常依赖单向工具。为此,本文提出Sanvaad——一个轻量级多模态无障碍框架,支持实时双向沟通。针对聋人用户,系统采用MediaPipe手部关键点构建手语识别模块,因其高效低耗,可在边缘设备上流畅运行,无需专用硬件。通过语音转手语组件,手机语音可被解析为预设短语,并生成对应GIF或字母可视化内容。针对视障者,框架提供无屏幕语音接口,集成多语言语音识别、文本摘要与文本转语音功能。所有组件通过Streamlit界面整合,兼容桌面与移动端。Sanvaad旨在通过统一框架结合轻量计算机视觉与语音处理技术,为包容性沟通提供实用路径。
原文摘要 · Abstract (English)
Communication between deaf users, visually im paired users, and the general hearing population often relies on tools that support only one direction of interaction. To address this limitation, this work presents Sanvaad, a lightweight multimodal accessibility framework designed to support real time, two-way communication. For deaf users, Sanvaad includes an ISL recognition module built on MediaPipe landmarks. MediaPipe is chosen primarily for its efficiency and low computational load, enabling the system to run smoothly on edge devices without requiring dedicated hardware. Spoken input from a phone can also be translated into sign representations through a voice-to-sign component that maps detected speech to predefined phrases and produces corresponding GIFs or alphabet-based visualizations. For visually impaired users, the framework provides a screen free voice interface that integrates multilingual speech recognition, text summarization, and text-to-speech generation. These components work together through a Streamlit-based interface, making the system usable on both desktop and mobile environments. Overall, Sanvaad aims to offer a practical and accessible pathway for inclusive communication by combining lightweight computer vision and speech processing tools within a unified framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。