为智能眼镜设计阅读识别系统,自动判断用户何时在阅读。
Reading Recognition in the Wild
- 用多模态数据(视觉、视线、头部姿态)识别阅读行为。
- 构建首个大规模野外阅读数据集,含100小时真实场景视频。
- 模型可灵活组合模态,适用于多样真实使用场景。
为实现始终在线的智能眼镜中的第一人称情境智能,必须记录用户与世界的交互,包括阅读行为。本文提出一项新任务——阅读识别,用于判断用户是否在阅读。我们首次构建了大规模多模态阅读数据集Reading in the Wild,包含100小时在多样化、真实场景下的阅读与非阅读视频。我们识别出三种有效模态:第一人称RGB图像、眼动追踪和头部姿态,并提出一个灵活的Transformer模型,可单独或联合使用这些模态完成任务。实验表明,这些模态对任务具有相关性和互补性,并验证了高效编码各模态的方法。此外,该数据集有助于分类阅读类型,将现有受控环境下的阅读理解研究扩展至更大规模、更多样性和更真实的场景。
原文摘要 · Abstract (English)
To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of reading recognition to determine when the user is reading. We first introduce the first-of-its-kind large-scale multimodal Reading in the Wild dataset, containing 100 hours of reading and non-reading videos in diverse and realistic scenarios. We then identify three modalities (egocentric RGB, eye gaze, head pose) that can be used to solve the task, and present a flexible transformer model that performs the task using these modalities, either individually or combined. We show that these modalities are relevant and complementary to the task, and investigate how to efficiently and effectively encode each modality. Additionally, we show the usefulness of this dataset towards classifying types of reading, extending current reading understanding studies conducted in constrained settings to larger scale, diversity and realism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。