用眼球追踪聚焦事件相机数据,低光下高效读取文字
Reading in the Dark with Foveated Event Vision
- 根据用户眼动聚焦事件流,大幅压缩数据带宽
- 在低光场景中实现精准文本识别,带宽仅需RGB相机的1/2400
- 结合合成数据重建与多模态大模型,适合可穿戴设备
当前配备RGB摄像头的智能眼镜在低光照和高速运动场景下难以感知环境,因运动模糊及帧相机动态范围有限。此外,密集图像采集需高带宽与高功耗,导致电池快速耗尽。这些挑战对文本识别算法尤为突出。本文提出一种基于事件相机的新型光学字符识别(OCR)方法,利用用户眼动信息对事件流进行聚焦处理,使带宽降低约98%。通过合成数据训练深度二值重建,并结合多模态大语言模型实现OCR,性能优于传统方案。实验表明,该方法可在RGB相机失效的低光环境中读取文字,带宽消耗仅为可穿戴RGB相机的1/2400。
原文摘要 · Abstract (English)
Current smart glasses equipped with RGB cameras struggle to perceive the environment in low-light and high-speed motion scenarios due to motion blur and the limited dynamic range of frame cameras. Additionally, capturing dense images with a frame camera requires large bandwidth and power consumption, consequently draining the battery faster. These challenges are especially relevant for developing algorithms that can read text from images. In this work, we propose a novel event-based Optical Character Recognition (OCR) approach for smart glasses. By using the eye gaze of the user, we foveate the event stream to significantly reduce bandwidth by around 98% while exploiting the benefits of event cameras in high-dynamic and fast scenes. Our proposed method performs deep binary reconstruction trained on synthetic data and leverages multimodal LLMs for OCR, outperforming traditional OCR solutions. Our results demonstrate the ability to read text in low light environments where RGB cameras struggle while using up to 2400 times less bandwidth than a wearable RGB camera.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。