构建多语言沉浸式教育AI平台,实现语音、翻译与手语实时转换。
AI-Driven Modular Services for Accessible Multilingual Education in Immersive Extended Reality Settings: Integrating Speech Processing, Translation, and Sign Language Rendering
- 模块化整合6项AI服务,支持语音识别到手语动画的全流程处理。
- AWS Polly延迟最低,EuroLLM 1.7B在翻译任务中表现优于NLLB。
- 适用于无障碍多语言教学,契合欧盟数字包容目标。
本文提出一个模块化平台,集成六项AI服务:通过OpenAI Whisper进行自动语音识别,使用Meta NLLB实现多语言翻译,借助AWS Polly进行语音合成,利用RoBERTa进行情感分类,采用flan t5 base samsum进行对话摘要,以及通过Google MediaPipe实现国际手语(IS)渲染。基于一段IS手势记录语料库,提取手部关键点坐标,并映射至虚拟现实(VR)环境中的三维角色动画。技术验证包括各AI组件的基准测试,涵盖语音合成服务与多语言翻译模型(NLLB 200与EuroLLM 1.7B)的对比评估。结果表明该平台适合实时扩展现实(XR)部署。语音合成测试显示AWS Polly延迟最低且性价比高;EuroLLM 1.7B Instruct变体在翻译任务中获得更高BLEU分数,优于NLLB。这些发现证明在XR环境中协同跨模态AI服务实现可访问性多语言教学的可行性。模块化设计支持独立扩展与适配不同教育场景,为符合欧盟数字无障碍目标的公平学习解决方案奠定基础。
原文摘要 · Abstract (English)
This work introduces a modular platform that brings together six AI services, automatic speech recognition via OpenAI Whisper, multilingual translation through Meta NLLB, speech synthesis using AWS Polly, emotion classification with RoBERTa, dialogue summarisation via flan t5 base samsum, and International Sign (IS) rendering through Google MediaPipe. A corpus of IS gesture recordings was processed to derive hand landmark coordinates, which were subsequently mapped onto three dimensional avatar animations inside a virtual reality (VR) environment. Validation comprised technical benchmarking of each AI component, including comparative assessments of speech synthesis providers and multilingual translation models (NLLB 200 and EuroLLM 1.7B variants). Technical evaluations confirmed the suitability of the platform for real time XR deployment. Speech synthesis benchmarking established that AWS Polly delivers the lowest latency at a competitive price point. The EuroLLM 1.7B Instruct variant attained a higher BLEU score, surpassing NLLB. These findings establish the viability of orchestrating cross modal AI services within XR settings for accessible, multilingual language instruction. The modular design permits independent scaling and adaptation to varied educational contexts, providing a foundation for equitable learning solutions aligned with European Union digital accessibility goals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。