用合成数据训练大模型,突破手写公式识别的数据瓶颈。
Towards Scalable Training for Handwritten Mathematical Expression Recognition
- 构建可扩展数据引擎,生成8000万条高质量LaTeX公式数据。
- 训练出首个大规模手写公式识别模型TexTeller,性能达当前最优。
- 开源全部数据与代码,推动领域研究进展。
大型基础模型通过在海量数据上进行可扩展训练取得了显著性能提升。然而,手写数学表达式识别(HMER)因数据稀缺而受限,主要由于人工标注过程繁重且成本高昂。为此,我们提出一种新方法,将有限的手写公式与大规模LaTeX渲染公式相结合,开发了一个可扩展的数据引擎,用于生成复杂且一致的LaTeX序列。基于该引擎,我们构建了迄今最大的公式数据集Tex80M,包含超过8000万条高质量训练样本。随后,我们提出首个大规模训练的HMER模型TexTeller,通过混合训练Tex80M与较小规模的真实手写数据集实现。庞大的训练数据和优化的训练流程使TexTeller在几乎所有基准测试中达到当前最优(SOTA)性能。为推进该领域发展,我们将公开发布完整模型、整个数据集及全部代码库,支持后续研究建立在我们的成果之上。
原文摘要 · Abstract (English)
Large foundation models have achieved significant performance gains through scalable training on massive datasets. However, the field of \textbf{H}andwritten \textbf{M}athematical \textbf{E}xpression \textbf{R}ecognition (HMER) has been impeded by the scarcity of data, primarily due to the arduous and costly process of manual annotation. To bridge this gap, we propose a novel method integrating limited handwritten formulas with large-scale LaTeX-rendered formulas by developing a scalable data engine to generate complex and consistent LaTeX sequences. With this engine, we built the largest formula dataset to date, termed \texttt{Tex80M}, comprising over 80 million high-quality training instances. Then we propose \texttt{TexTeller}, the first HMER model trained at scale, by mix-training \texttt{Tex80M} with a relatively small HME dataset. The expansive training dataset and our refined pipeline have equipped \texttt{TexTeller} with state-of-the-art (SOTA) performance across nearly all benchmarks. To advance the field, we will openly release our complete model, entire dataset, and full codebase, enabling further research building upon our contributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。