arXiv:2608.22183cs.CVcs.IR2026-08

用多个模型一致判断,比像素级重渲染更准地识别化学结构图

VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR

论文配图:VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR
图 1 · 摘自论文原文
  • 通过多个不同架构模型的预测一致性判断可靠性
  • 一致判断准确率高达91.6%,远超像素重渲染的54.7%
  • 适合构建无真实标签的化学结构数据集,尤其适用于真实文献图像

光学化学结构识别(OCSR)将文献中的二维分子图转化为SMILES,对构建大规模化学训练数据至关重要。在缺乏真实标签的情况下,需自动识别不可靠预测。在263张经验证的ACS期刊图像上,对比了模型置信度、重渲染相似性及多模型一致性三种无标签信号。像素空间重渲染表现仅略优于随机(AUROC 0.547,95% CI [0.465,0.629]),最优阈值使每图正确标签数从0.745降至0.205。而四个不同架构模型的一致性达到AUROC 0.916([0.880,0.952])。两票通过规则可接受81.7%图像,精度88.8%;三票通过规则接受52.1%,精度达98.5%。该规律在CLEF-IP、UOB和USPTO数据集上同样成立。合成基准因重渲染自然贴近输入,掩盖了此差异。引入物质过滤器后,剔除2193个由通配符和R基团引发的错误一致性,三票规则成功排除全部68个通用图。将VERDICT应用于PMC开放获取文献,获得4833个分子的6146个结构标签;两名独立样本中400个标签经化学家人工审核,三票层级精度达0.995,两票层级为0.958。VERDICT因此可实现多模态分子数据库的可靠标注,连接结构图像、机器可读表示与原始文献信息。在SES AI的Molecular Universe平台中,还作为基于图像的分子检索接口。

原文摘要 · Abstract (English)

Optical Chemical Structure Recognition (OCSR) converts 2D molecular depictions in the published literature into SMILES, and is increasingly important for constructing large-scale chemical training datasets. Automation at that scale requires identifying unreliable predictions in the absence of ground truth. Three families of label-free signals were compared on $263$ ACS journal depictions with verified ground truth: model confidence, re-rendering similarity, and agreement among recognizers. Pixel-space re-rendering performed little better than chance (AUROC $0.547$, $95\%$ CI $[0.465,0.629]$), and an oracle-tuned threshold on it reduced correct labels per image from $0.745$ to $0.205$. Agreement among four architecturally distinct recognizers instead reached an AUROC of $0.916$ ($[0.880,0.952]$). The two-of-four rule accepted $81.7\%$ of images at $88.8\%$ precision, the three-of-four rule $52.1\%$ at $98.5\%$. The same pattern held on CLEF-IP, UOB, and USPTO. This distinction is obscured on synthetic benchmarks, where re-rendered predictions naturally resemble their inputs. A substance filter removed $2{,}193$ false agreements on wildcards and R-group fragments, after which the three-of-four rule rejected all $68$ generic depictions. VERDICT was then applied to PMC Open Access, producing $6{,}146$ structure labels for $4{,}833$ molecules; chemist adjudication of $400$ released labels in two independent samples yielded precisions of $0.995$ for the three-of-four tier and $0.958$ for the two-of-four tier. VERDICT therefore enables validated labels for multimodal molecular databases linking structure images, machine-readable representations, and source-publication information. In SES AI's Molecular Universe platform, VERDICT further serves as an image-based interface for searching and retrieving molecular records.

化学结构识别多模型一致性数据标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。