用真实化学结构图训练,大幅提升模型在真实文档中的识别准确率
Real Data Closes Synthetic-to-Real Gap in Optical Chemical Structure Recognition

- 用真实专利、期刊图等数据混合训练,显著提升模型性能
- 在真实数据上准确率从16%提升至最高84%(如USPTO)
- 不同基础模型对真实数据敏感度不同,需协同选择
数百万化学结构仅以图像形式出现在专利和论文中,大规模利用需读图。尽管合成图像上的光学化学结构识别(OCSR)已接近解决,但在真实文档上仍困难:基线模型Qwen2.5-VL-7B在合成图像上准确率超91%,但在三个真实基准(ACS、CLEF-IP、USPTO)上低于16%。为探究改进关键,21个识别器在合成结构与标注的真实图像(专利、期刊图、手绘集)混合数据上微调,调整视觉语言模型(VLM)基座、真实数据占比及视觉塔适配策略。标注真实数据带来最大提升:对Qwen2.5-VL,ACS精确匹配从0.15升至0.37(9.5%真实数据)和0.46(50.2%真实数据);三模型对照实验复现此趋势。视觉塔LoRA对Qwen无效(+0.00,配对p=1.00),显著帮助InternVL3-8B(+22.8~+34.6 pt),适度帮助GLM-4.1V-9B(+1.0~+9.6 pt),表明其效果依赖基座模型。最优配置在干净渲染图上达0.96精确匹配,在ACS、CLEF-IP、UOB、USPTO上分别达0.49、0.65、0.84、0.76。无真实数据时基座模型差距最大(0.21),70%真实数据下缩至0.06并重排排名,说明基座模型与真实数据混合需协同选择。小规模实验显示手写图像转LaTeX、图表转表格任务中基座模型排名亦变化。总体而言,视觉结构识别的模型与适配选择应基于目标任务评估。
原文摘要 · Abstract (English)
Millions of chemical structures appear in patents and papers only as drawings, and using that information at scale requires reading the drawings. OCSR appears nearly solved on synthetic images yet remains difficult on real documents: the starting recognizer, Qwen2.5-VL-7B, exceeds 91% accuracy on synthetic renders but falls below 16% on three real-world benchmarks (ACS, CLEF-IP, USPTO). To identify the main source of improvement, 21 recognizers were fine-tuned on mixtures of synthetically rendered structures and labeled real depictions from patents, journal figures, and hand-drawn collections, varying the vision language model (VLM) base, the fraction of real training data, and the vision-tower adaptation strategy. Labeled real training images make the largest difference. For Qwen2.5-VL, ACS exact match rises from 0.15 with no real data to 0.37 at 9.5% and 0.46 at 50.2%; a controlled experiment across three base models reproduces the trend. A vision-tower LoRA, in contrast, does nothing for Qwen (+0.00, paired p=1.00), substantially helps InternVL3-8B (+22.8 to +34.6 pt), and modestly helps GLM-4.1V-9B (+1.0 to +9.6 pt), so its value depends on the base model. The best configuration reaches 0.96 exact match on clean renders and 0.49, 0.65, 0.84, and 0.76 on ACS, CLEF-IP, UOB, and USPTO, respectively. Gaps between base models are largest without real data (0.21), shrink to 0.06 at 70% real data, and reorder the ranking; base model and real-data mixture must therefore be selected together. Small-scale experiments on handwritten image-to-LaTeX recognition and chart-to-table conversion show that base-model rankings also vary beyond chemistry. More generally, model and adaptation choices for visual structure recognition should be evaluated on the target task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。