用双自洽强化学习让AI精准生成可执行的科学绘图代码
Scientific Graphics Program Synthesis via Dual Self-Consistency Reinforcement Learning
- 提出双自洽强化学习,通过代码往返验证提升生成质量
- 构建23万条可执行数据集与多维度评测基准,覆盖11个学科
- 模型在视觉与结构上均超越谷歌、通义等大模型
图形程序合成对解读和编辑可视化数据至关重要,能将静态图像还原为可编辑的TikZ代码。尽管TikZ是科学图表的标准工具,其对空间精度的严格要求给多模态大模型带来挑战。当前主要存在两大瓶颈:(1) 数据质量不足,现有图像-TikZ语料库常缺乏可执行性与可靠视觉对齐;(2) 缺乏评估基准,难以同时衡量结构与视觉保真度。为此,我们提出闭环框架:SciTikZ-230K,一个由执行中心数据引擎构建的大规模高质量数据集,涵盖11个科学领域;SciTikZ-Bench,覆盖基础几何到复杂层级结构的多维评测基准。进一步引入双自洽强化学习优化范式,利用往返验证惩罚退化代码,提升整体自洽性。基于此,训练出的SciTikZer-8B模型表现领先,持续优于Gemini-2.5-Pro与Qwen3-VL-235B-A22B-Instruct等主流模型。
原文摘要 · Abstract (English)
Graphics Program Synthesis is pivotal for interpreting and editing visual data, effectively facilitating the reverse-engineering of static visuals into editable TikZ code. While TikZ is the de facto standard for scientific schematics due to its programmatic flexibility, its requirement for rigorous spatial precision presents a significant challenge for Multimodal Large Language Models. Progress is currently stifled by two primary gaps: (1) Data Quality Gap: existing image-TikZ corpora often lack strict executability and reliable visual alignment; (2) Evaluation Gap: a lack of benchmarks for both structural and visual fidelity. To address these, we present a closed-loop framework featuring: SciTikZ-230K, a large-scale, high-quality dataset from our Execution-Centric Data Engine covering 11 diverse scientific disciplines; SciTikZ-Bench, a multifaceted benchmark spanning from basic geometric constructs to intricate hierarchical schematics to evaluate both visual fidelity and structural logic. To further broaden the scope of visual-code optimization methodology, we introduce a novel Dual Self-Consistency Reinforcement Learning optimization paradigm, which utilizes Round-Trip Verification to penalize degenerate code and boost overall self-consistency. Empowered by these, our trained model SciTikZer-8B achieves state-of-the-art performance, consistently outperforming proprietary giants like Gemini-2.5-Pro and massive models like Qwen3-VL-235B-A22B-Instruct.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。