手机手握拍文档视频,自动识别多页并拼接
Handheld Video Document Scanning: A Robust On-Device Model for Multi-Page Document Scanning
- 用深度学习模型从手持拍摄的视频中自动检测页面翻转
- 在PUCIT数据集上达到当前最优性能,准确识别翻页时刻
- 适合移动端实时使用,无需三脚架等外设
智能手机文档拍摄应用已成为数字化文件的常用工具。尽管质量低于专用扫描仪,但对多数用户而言,用手机拍摄更便捷。然而,当处理多页文档时,手动逐页拍摄极为耗时。本文提出一种新方法,通过用户手持手机拍摄文档翻页视频,自动识别并分割多页内容。与以往需固定设备(如三脚架)的方法不同,本方法专为手持不稳定场景设计,训练时增强对运动模糊和抖动的鲁棒性。主要贡献包括:(1) 一个高效、精准且适用于移动端的深度学习模型;(2) 针对视频文档扫描的新型数据采集与标注方法;(3) 在PUCIT page turn数据集上取得当前最优结果。
原文摘要 · Abstract (English)
Document capture applications on smartphones have emerged as popular tools for digitizing documents. For many individuals, capturing documents with their smartphones is more convenient than using dedicated photocopiers or scanners, even if the quality of digitization is lower. However, using a smartphone for digitization can become excessively time-consuming and tedious when a user needs to digitize a document with multiple pages. In this work, we propose a novel approach to automatically scan multi-page documents from a video stream as the user turns through the pages of the document. Unlike previous methods that required constrained settings such as mounting the phone on a tripod, our technique is designed to allow the user to hold the phone in their hand. Our technique is trained to be robust to the motion and instability inherent in handheld scanning. Our primary contributions in this work include: (1) an efficient, on-device deep learning model that is accurate and robust for handheld scanning, (2) a novel data collection and annotation technique for video document scanning, and (3) state-of-the-art results on the PUCIT page turn dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。