为视障儿童打造的离线语音学习助手,能实时识别学习困难并自适应调整。
Kutti AI: A Voice-First, Offline-Capable Learning Companion with Real-Time Struggle Detection for Visually-Impaired Children
- 用语音交互替代视觉界面,支持听、说、反馈全链路对话学习。
- 通过延迟、错误次数和停顿分析,实时判断孩子是否卡住并及时提示。
- 支持多语言混用和口音差异,可在无网络环境下运行,适合资源匮乏地区。
多数儿童教育科技依赖视觉界面,排除了全球约140万盲童及更多低视力儿童。本文提出Kutti AI,一种以语音为核心的离线学习伴侣:儿童通过口语对话学习课程内容,口头作答并接收语音反馈,全程无需视觉元素。系统在普通移动设备上实现三项实用机制:(1)多信号挣扎检测引擎,结合响应延迟分析、错误尝试追踪与关键词停顿检测,实时判断是否需提供提示或简化问题;(2)多层跨语言答案匹配管道,融合语言感知翻译/转写、基于Levenshtein的距离模糊匹配与文本归一化,避免因语言切换或发音差异被扣分;(3)基于本地自动语音识别(ASR)模型的离线语音处理流程,适用于低网络覆盖地区。我们描述了系统架构、交互流程与无障碍设计决策,并报告了支持英语与泰米尔语的黑客松原型的定性观察。讨论了经验教训,并规划向目标用户进行正式评估。Kutti AI展示了一个小巧而精心设计的语音优先系统如何降低早期教育的可及性与经济门槛。
原文摘要 · Abstract (English)
Most educational technology for children is built around visual interfaces, which excludes the many children worldwide who live with visual impairment -- an estimated 1.4 million children are blind and many more have low vision. We present Kutti AI, a voice-first learning companion designed so that audio is the primary and sufficient interface: children learn curriculum concepts through spoken conversation, respond by speaking, and receive spoken feedback, with no reliance on visual elements. The system contributes three practical mechanisms for accessible, adaptive learning on commodity mobile hardware: (1) a multi-signal struggle-detection engine that combines response-latency analysis, wrong-attempt tracking, and keyword-based hesitation detection to decide, in real time, when to offer hints or simplify a question; (2) a multi-layered cross-language answer-matching pipeline that combines language-aware translation/transliteration, Levenshtein-based fuzzy matching, and text normalization so that children are not penalized for code-switching or pronunciation variation; and (3) an offline-first speech pipeline using an on-device automatic speech recognition (ASR) model, enabling use in low-connectivity settings common in underserved communities. We describe the architecture, the interaction flow, and the design decisions that prioritize accessibility, and we report qualitative observations from a hackathon prototype supporting English and Tamil. We discuss lessons learned and outline a path toward formal evaluation with target users. Kutti AI illustrates how a small, carefully-engineered voice-first system can lower both accessibility and financial barriers to early education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。