科学图表编辑不能只改像素,需理解数据结构。
Charts Are Not Images: On the Challenges of Scientific Chart Editing
- 提出图编辑新基准FigEdit,覆盖3万+样本和10类图表
- 现有模型在复杂编辑任务中表现差,多因忽略数据结构
- 适合研究视觉与语义融合的图表生成与编辑者
生成模型在自然图像编辑中表现优异,但将其应用于科学图表存在根本误解:图表不仅是像素排列,更是受图形语法规则约束的结构化数据可视化。因此,图表编辑本质是结构化变换问题,而非像素操作。为解决这一错位,我们提出 extit{FigEdit}——一个大规模科学图表编辑基准,包含超30,000个样本,涵盖10种图表类型及丰富复杂的编辑指令。该基准分为五类逐步提升难度的任务:单次编辑、多次编辑、对话式编辑、基于视觉引导的编辑和风格迁移。对多种前沿模型的评估显示,它们在科学图表上表现不佳,普遍无法正确执行所需的结构化变换。分析还表明,传统指标(如SSIM、PSNR)难以衡量图表编辑的语义正确性。本基准揭示了像素级操作的深层局限,并为未来结构感知模型的发展提供了坚实基础。通过开源 extit{FigEdit} (https://github.com/adobe-research/figure-editing),我们旨在推动结构感知图表编辑的系统性进展,建立公平比较的共同标准,激励下一代兼具视觉与语义理解能力的模型研究。
原文摘要 · Abstract (English)
Generative models, such as diffusion and autoregressive approaches, have demonstrated impressive capabilities in editing natural images. However, applying these tools to scientific charts rests on a flawed assumption: a chart is not merely an arrangement of pixels but a visual representation of structured data governed by a graphical grammar. Consequently, chart editing is not a pixel-manipulation task but a structured transformation problem. To address this fundamental mismatch, we introduce \textit{FigEdit}, a large-scale benchmark for scientific figure editing comprising over 30,000 samples. Grounded in real-world data, our benchmark is distinguished by its diversity, covering 10 distinct chart types and a rich vocabulary of complex editing instructions. The benchmark is organized into five distinct and progressively challenging tasks: single edits, multi edits, conversational edits, visual-guidance-based edits, and style transfer. Our evaluation of a range of state-of-the-art models on this benchmark reveals their poor performance on scientific figures, as they consistently fail to handle the underlying structured transformations required for valid edits. Furthermore, our analysis indicates that traditional evaluation metrics (e.g., SSIM, PSNR) have limitations in capturing the semantic correctness of chart edits. Our benchmark demonstrates the profound limitations of pixel-level manipulation and provides a robust foundation for developing and evaluating future structure-aware models. By releasing \textit{FigEdit} (https://github.com/adobe-research/figure-editing), we aim to enable systematic progress in structure-aware figure editing, provide a common ground for fair comparison, and encourage future research on models that understand both the visual and semantic layers of scientific charts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。