arXiv:2501.11233cs.IRcs.CL2025-01中稿 · ECIR 2025被引 4

用自然语言直接编辑PDF中的图表,无需原始数据或代码。

PlotEdit: Natural Language-Driven Accessible Chart Editing in PDFs via Multimodal LLM Agents

  • 五位AI代理协作:提取数据、识别风格、找渲染代码、拆解指令、实现修改。
  • 在ChartCraft数据集上各项编辑任务表现优于现有方法,视觉保真度高。
  • 适合需要快速修改图表的科研人员、视障用户及初学者使用。

图表可视化虽对数据解读与传播至关重要,但通常以图像形式存在于PDF中,缺乏源数据表和样式信息。为实现对PDF或数字扫描图中图表的有效编辑,我们提出PlotEdit,一种基于自省式多模态大模型代理的新型多代理框架,支持自然语言驱动的端到端图表图像编辑。PlotEdit协调五个大模型代理:(1) Chart2Table用于数据表提取,(2) Chart2Vision用于风格属性识别,(3) Chart2Code用于检索渲染代码,(4) 指令分解代理将用户请求解析为可执行步骤,(5) 多模态编辑代理实现细微的图表组件修改——所有操作通过多模态反馈协调,确保视觉一致性。在ChartCraft数据集上,PlotEdit在风格、布局、格式及数据相关编辑任务中均优于现有基线方法,显著提升视障用户的可访问性,并提高新手编辑效率。

原文摘要 · Abstract (English)

Chart visualizations, while essential for data interpretation and communication, are predominantly accessible only as images in PDFs, lacking source data tables and stylistic information. To enable effective editing of charts in PDFs or digital scans, we present PlotEdit, a novel multi-agent framework for natural language-driven end-to-end chart image editing via self-reflective LLM agents. PlotEdit orchestrates five LLM agents: (1) Chart2Table for data table extraction, (2) Chart2Vision for style attribute identification, (3) Chart2Code for retrieving rendering code, (4) Instruction Decomposition Agent for parsing user requests into executable steps, and (5) Multimodal Editing Agent for implementing nuanced chart component modifications - all coordinated through multimodal feedback to maintain visual fidelity. PlotEdit outperforms existing baselines on the ChartCraft dataset across style, layout, format, and data-centric edits, enhancing accessibility for visually challenged users and improving novice productivity.

图表编辑多模态LLMPDF处理无障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。