用普通摄像头实现眼控键盘,支持同步异步输入,适合资源有限的残障人士使用。
Multimodal Appearance based Gaze-Controlled Virtual Keyboard with Synchronous Asynchronous Interaction for Low-Resource Settings
- 结合深度学习与普通摄像头,设计双模式眼控键盘
- 同步模式下打字速度达10.94字/分钟,信息传输率63.56比特/分钟
- 无需专用设备,适合低资源环境下的无障碍沟通
过去十年,行动与语言障碍者对通信设备的需求上升。眼动追踪成为免手操作通信的有前景方案,但传统基于外观的界面常面临准确率低、不自主眼动及复杂指令集难用等问题。本文提出一种多模态外观眼控虚拟键盘,利用深度学习与标准摄像头硬件,支持同步与异步两种指令选择模式。该键盘提供九项菜单式指令,可输入包括大小写字母、标点符号和删除功能在内的56个英文字符。在20名健康参与者中,通过鼠标、眼动仪和未改装网络摄像头三种方式完成指定打字任务。结果显示:同步模式下,摄像头输入平均打字速度为10.94±1.89字/分钟,信息传输率(ITR)为63.56±11比特/分钟(字母级),命令级为80.29±15.72比特/分钟。系统在摄像头输入下表现出良好可用性与低认知负荷,体现了以用户为中心的设计,具有在低资源环境中推广的潜力。
原文摘要 · Abstract (English)
Over the past decade, the demand for communication devices has increased among individuals with mobility and speech impairments. Eye-gaze tracking has emerged as a promising solution for hands-free communication; however, traditional appearance-based interfaces often face challenges such as accuracy issues, involuntary eye movements, and difficulties with extensive command sets. This work presents a multimodal appearance-based gaze-controlled virtual keyboard that utilises deep learning in conjunction with standard camera hardware, incorporating both synchronous and asynchronous modes for command selection. The virtual keyboard application supports menu-based selection with nine commands, enabling users to spell and type up to 56 English characters, including uppercase and lowercase letters, punctuation, and a delete function for corrections. The proposed system was evaluated with twenty able-bodied participants who completed specially designed typing tasks using three input modalities: (i) a mouse, (ii) an eye-tracker, and (iii) an unmodified webcam. Typing performance was measured in terms of speed and information transfer rate (ITR) at both command and letter levels. Average typing speeds were 18.3+-5.31 letters/min (mouse), 12.60+-2.99letters/min (eye-tracker, synchronous), 10.94 +- 1.89 letters/min (webcam, synchronous), 11.15 +- 2.90 letters/min (eye-tracker, asynchronous), and 7.86 +- 1.69 letters/min (webcam, asynchronous). ITRs were approximately 80.29 +- 15.72 bits/min (command level) and 63.56 +- 11 bits/min (letter level) with webcam in synchronous mode. The system demonstrated good usability and low workload with webcam input, highlighting its user-centred design and promise as an accessible communication tool in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。