构建可编辑的科学图表代码生成基准,提升论文图示复现性。
SciFigure2Code: An AI-Reconstructed Benchmark for Scientific Figure-to-Code

- 用AI重建图表布局与视觉层级,生成可运行的Python代码。
- 包含6740个审核图表和337个跨领域测试集,覆盖31种图表类型。
- 适合研究可解释性、自动化科研工具与图表复现的开发者使用。
科学图表是研究结论呈现与复用的关键界面,但最终发表的图表通常不公开其数据或绘图代码。从像素中恢复原始数据与代码存在严重不确定性。本文提出SciFigure2Code,一个基于AI重构的基准,专注于可视化呈现的复现:生成可编辑的Python程序,保留图表的版式、阅读顺序、几何结构、视觉层次、编码方式、标注和排版。通过角色分工的Codex智能体生成、执行、视觉优化与审计,产出银标准的呈现代码。该协议将发布图表转化为可审计的参考包,包含6740个经审核的图表及SciFigureBench测试集——337个面板,涵盖31种图表子类型、5个学科领域与3种复杂度等级。在14个零样本模型上测试,图像仅与带标题辅助的设置下,执行、多组件布局、坐标轴、图例与科学标签仍表现较弱。Claude Opus 4.7在纯图像输入中表现最佳,Claude Opus 4.6在标题辅助下领先,两阶段‘规划-生成’提示策略显著提升所有四款模型的整体表现。SciFigure2Code为构建可编辑、视觉忠实的科学图表生成代理提供了可审计的测试平台。
原文摘要 · Abstract (English)
Scientific figures are the interface through which research claims are inspected and reused, but final published panels rarely expose the data or plotting code that produced them. Recovering this hidden provenance from pixels is therefore underdetermined. We introduce SciFigure2Code, an AI-reconstructed benchmark that instead evaluates presentation recovery: generating editable Python programs that preserve how a scientific panel is arranged and read. Role-specialized Codex agents generate, execute, visually refine, and audit silver-standard presentation programs that capture geometry, visual hierarchy, encodings, annotations, and typography without claiming to recover original measurements or author source code. This reconstruction-and-audit protocol turns final published panels into auditable reference packages; the resulting resource contains 6,740 reviewed panels and SciFigureBench, a balanced 337-panel test set across 31 chart subtypes, five domains, and three complexity levels. Across 14 zero-shot models in image-only and caption-assisted settings, execution, multi-component layouts, axes, legends, and scientific labels remain weak. Claude Opus 4.7 achieves the highest image-only Overall score, Claude Opus 4.6 leads caption-assisted reconstruction, and two-stage plan-then-code prompting improves Overall for all four tested models. SciFigure2Code provides an auditable testbed for agents that construct editable, visually faithful scientific figure presentations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。