InkFM可一键识别28种文字与手绘内容,提升数字笔记理解效率。
InkFM: A Foundational Model for Full-Page Online Handwritten Note Understanding
- 基于多任务训练的统一模型,兼顾文字、公式与页面元素分割。
- 在文本行分割上超越docTR,手写文字识别达最新水平。
- 适合开发手写输入应用,支持快速微调适配各类数据集。
平板与触控笔在笔记记录中日益普及。为优化体验并实现高效工作流,准确解析数字手写笔记内容至关重要。我们提出一种名为InkFM的基础模型,用于分析整页手写内容。该模型通过多样化任务训练,具备独特能力:支持28种不同文字的文本识别、数学表达式识别,并能将页面分割为文本、绘图等独立元素。实验表明,这些任务可在单一模型中有效统一,实现优于公开基线(如docTR)的文本行分割效果。对公共数据集进行微调或LoRA微调后,模型在文本识别(DeepWriting、CASIA、SCUT、Mathwriting数据集)和草图分类(QuickDraw)方面达到当前最优性能。InkFM的可扩展性为开发手写输入应用提供了强大起点。
原文摘要 · Abstract (English)
Tablets and styluses are increasingly popular for taking notes. To optimize this experience and ensure a smooth and efficient workflow, it's important to develop methods for accurately interpreting and understanding the content of handwritten digital notes. We introduce a foundational model called InkFM for analyzing full pages of handwritten content. Trained on a diverse mixture of tasks, this model offers a unique combination of capabilities: recognizing text in 28 different scripts, mathematical expressions recognition, and segmenting pages into distinct elements like text and drawings. Our results demonstrate that these tasks can be effectively unified within a single model, achieving SoTA text line segmentation out-of-the-box quality surpassing public baselines like docTR. Fine- or LoRA-tuning our base model on public datasets further improves the quality of page segmentation, achieves state-of the art text recognition (DeepWriting, CASIA, SCUT, and Mathwriting datasets) and sketch classification (QuickDraw). This adaptability of InkFM provides a powerful starting point for developing applications with handwritten input.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。