用图像识别+自然语言生成,让渐冻症患者高效打字。
ImageTalk: Designing a Multimodal AAC Text Generation System Driven by Image Recognition and Natural Language Generation
- 通过图像识别自动转文本,减少输入步骤
- 键入次数减少95.6%,沟通效率大幅提升
- 适合渐冻症患者及辅助沟通系统研究者
患有运动神经元疾病(plwMND)的人群常因言语和运动障碍,需依赖增强与替代通信(AAC)系统。传统基于符号的AAC系统词汇量有限,而文本输入方式则沟通速率低。为帮助plwMND更高效表达需求,我们通过代理用户与真实用户参与的设计阶段,迭代开发了名为ImageTalk的多模态文本生成系统。该系统在实验中实现95.6%的键入节省,保持稳定性能并获得高用户满意度。我们提炼出三条面向AI辅助文本生成系统的设计准则,并提出四类适用于AAC场景的用户需求层级,为该领域未来研究提供指导。
原文摘要 · Abstract (English)
People living with Motor Neuron Disease (plwMND) frequently encounter speech and motor impairments that necessitate a reliance on augmentative and alternative communication (AAC) systems. This paper tackles the main challenge that traditional symbol-based AAC systems offer a limited vocabulary, while text entry solutions tend to exhibit low communication rates. To help plwMND articulate their needs about the system efficiently and effectively, we iteratively design and develop a novel multimodal text generation system called ImageTalk through a tailored proxy-user-based and an end-user-based design phase. The system demonstrates pronounced keystroke savings of 95.6%, coupled with consistent performance and high user satisfaction. We distill three design guidelines for AI-assisted text generation systems design and outline four user requirement levels tailored for AAC purposes, guiding future research in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。