用AI实时翻译手语并识别情绪,让聋人无障碍参与视频会议。
INTERACT: An AI-Driven Extended Reality Framework for Accesible Communication Featuring Real-Time Sign Language Interpretation and Emotion Recognition
- 构建基于XR的AI平台,通过3D化身实时呈现国际手语。
- 手语识别准确率90%,语音转文字正确率超85%,用户满意度92%。
- 适合教育、职场等场景,特别为聋人及多语言用户设计。
视频会议已成为专业协作的核心,但多数平台对聋人、重听者及多语言用户支持有限。世界卫生组织估计,全球超过4.3亿人需听力康复,这一数字预计到2050年将超过7亿。传统无障碍措施受限于高成本、资源匮乏和后勤障碍,而扩展现实(XR)技术为沉浸式、包容性沟通提供了新可能。本文提出INTERACT(包容性网络翻译与具身实时增强通信工具),一个基于AI的XR平台,集成实时语音转文字、3D化身呈现国际手语(ISL)、多语言翻译及情绪识别功能。系统基于CORTEX2框架,部署于Meta Quest 3头显,采用Whisper进行语音识别,NLLB实现多语言翻译,RoBERTa用于情绪分类,Google MediaPipe提取手势。分两阶段开展试点评估:首先由学术与产业界技术专家参与,随后邀请聋人群体试用。结果显示用户满意度达92%,转录准确率超过85%,情绪检测精度达90%,整体体验平均评分4.6/5.0,90%参与者愿继续参与测试。结果表明该平台在教育、文化与职业场景中具有显著应用潜力。完整试点数据与实施细节已作为开放研究文章发表于《Open Research Europe》[Tantaroudas et al., 2026a]。
原文摘要 · Abstract (English)
Video conferencing has become central to professional collaboration, yet most platforms offer limited support for deaf, hard-of-hearing, and multilingual users. The World Health Organisation estimates that over 430 million people worldwide require rehabilitation for disabling hearing loss, a figure projected to exceed 700 million by 2050. Conventional accessibility measures remain constrained by high costs, limited availability, and logistical barriers, while Extended Reality (XR) technologies open new possibilities for immersive and inclusive communication. This paper presents INTERACT (Inclusive Networking for Translation and Embodied Real-Time Augmented Communication Tool), an AI-driven XR platform that integrates real-time speech-to-text conversion, International Sign Language (ISL) rendering through 3D avatars, multilingual translation, and emotion recognition within an immersive virtual environment. Built on the CORTEX2 framework and deployed on Meta Quest 3 headsets, INTERACT combines Whisper for speech recognition, NLLB for multilingual translation, RoBERTa for emotion classification, and Google MediaPipe for gesture extraction. Pilot evaluations were conducted in two phases, first with technical experts from academia and industry, and subsequently with members of the deaf community. The trials reported 92% user satisfaction, transcription accuracy above 85%, and 90% emotion-detection precision, with a mean overall experience rating of 4.6 out of 5.0 and 90% of participants willing to take part in further testing. The results highlight strong potential for advancing accessibility across educational, cultural, and professional settings. An extended version of this work, including full pilot data and implementation details, has been published as an Open Research Europe article [Tantaroudas et al., 2026a].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。