让AI从论文修订中学习科学图表编辑,支持自然语言指令操作。
SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

- 基于论文版本对比学习编辑意图,从真实修订中提取训练信号。
- 通过技能演化机制提升编辑准确率,在验证集上持续优化。
- 适合需要频繁修改科研图表的研究者与自动化工具开发者。
论文图表编辑是科研日常中耗时的环节:作者在修订稿件时需重命名组件、调整版面布局、更新视觉样式。然而,基于自然语言指令自动完成这一流程极具挑战,因为科学图表是复杂的视觉信息图,包含示意图、图表、照片、注释和箭头等多种元素,遵循严格的视觉语法以支撑特定论点。为此,我们提出SciDiagramEdit,一个基准数据集与技能演化框架,从arXiv版本历史中挖掘前后图对,每一对均源自作者真实的修订意图。为应对编辑指令的多样性,采用代理式学习中的技能演化机制:代理提案者通过多轮执行轨迹不断优化智能体的技能规范。结果表明,该技能在独立验证集上的编辑准确率持续提升,证明自然论文修订是指导图编辑任务的有效训练信号。
原文摘要 · Abstract (English)
Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rearrange panels, and restyle visuals as they revise their manuscripts. Automating this editing workflow under a natural-language instruction, however, is challenging, because a scientific figure is a dense infographic in which heterogeneous visual elements such as schematics, plots, photos, captions, and arrows are composed under a tight visual grammar to advance a specific argument. To address this, we present SciDiagramEdit, a benchmark and skill-evolution framework that learns from natural paper revisions and operates on the figure's editable vector source, where users can inspect and co-edit individual primitives alongside the agent. Our benchmark mines before/after figure pairs from arXiv version histories, each grounded in the authors' own revision intent. To accommodate the diversity of editing instructions, we adopt agentic learning via skill evolution: an agentic proposer continually refines the agent's skill specification from execution traces over multiple epochs. The resulting skill progressively lifts edit accuracy on a held-out validation set, providing evidence that natural paper revisions are an effective training signal for instruction-driven figure editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。