实测步行速度和相机位置对辅助视力的OCR识别率影响,发现越快越歪越不准。
Evaluating OCR Performance for Assistive Technology: Effects of Walking Speed, Camera Placement, and Camera Type
- 在1-7米距离、0-75度视角下测试静态与动态场景下的OCR表现
- 步行速度从0.8到1.8米/秒时识别准确率下降,视角越宽越差
- 手机主摄像头+肩挂位置效果最好,谷歌Vision识别最准
光学字符识别(OCR)广泛应用于盲人和低视力人群的辅助技术中。然而,现有评估多基于静态数据集,难以反映实际移动使用中的挑战。本研究系统评估了静态与动态条件下OCR的表现。静态测试涵盖1-7米距离及0-75度水平视角;动态测试则改变步行速度(0.8–1.8米/秒),比较头戴、肩挂、手持三种相机位置,并使用智能手机主摄与超广角镜头。评测四种OCR引擎:Google Vision、PaddleOCR 3.0、EasyOCR和Tesseract。结果表明,识别准确率随步行速度增加和视角扩大而下降。Google Vision总体准确率最高,PaddleOCR 3.0为最强开源替代方案。手机主摄像头表现最优,肩挂位置平均表现最佳,但头戴、肩挂、手持三者差异无统计显著性。
原文摘要 · Abstract (English)
Optical character recognition (OCR), a process that converts printed or handwritten text into machine-readable form, is widely used in assistive technology for people with blindness and low vision. Yet most evaluations rely on static datasets that do not reflect the challenges of mobile use. In this study, we systematically evaluated OCR performance under both static and dynamic conditions. Static tests measured detection range across distances of 1-7 meters and viewing angles of 0-75 degrees horizontally. Dynamic tests examined the impact of motion by varying walking speed from slow (0.8 m/s) to very fast (1.8 m/s) and compared three camera mounting positions: head-mounted, shoulder-mounted, and handheld. We evaluated both a smartphone and smart glasses, using the phone's main and ultra-wide cameras. Four OCR engines were benchmarked to assess accuracy at different distances and viewing angles: Google Vision, PaddleOCR 3.0, EasyOCR, and Tesseract. PaddleOCR 3.0 was then used to evaluate OCR performance under dynamic walking conditions. Accuracy was computed at the character-level using the Levenshtein ratio against manually defined ground truth. Results showed that recognition accuracy declined with increased walking speed and wider viewing angles. Google Vision achieved the highest overall accuracy, with PaddleOCR close behind as the strongest open-source alternative. Across devices, the phone's main camera achieved the highest accuracy, and a shoulder-mounted placement yielded the highest average among body positions; however, differences among shoulder, head, and hand were not statistically significant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。