让AI生成图表时更讲事实,解决图像编辑中的准确性难题
Factuality Matters: When Image Generation and Editing Meet Structured Visuals
- 构建130万对高质量结构化图像数据集,带思维链标注
- 训练统一模型,在编辑任务中显著提升事实准确性
- 推出新评测基准StructBench,支持细粒度真实度评估
当前视觉生成模型在自然图像创作上表现优异,但在生成或编辑图表、示意图和数学图形等结构化视觉内容时存在明显不足,这类任务要求构图规划、文本渲染与多模态推理以保证事实正确性。为此,我们首次系统性地开展该领域研究,涵盖数据构建、模型训练与评测基准。首先,基于可执行绘图程序构建包含130万对高质量结构化图像的数据集,并添加思维链(chain-of-thought)标注。在此基础上,训练一个融合视觉语言模型与FLUX.1 Kontext的统一模型,通过轻量级连接器增强多模态理解能力。采用三阶段训练流程实现特征对齐、知识注入与推理增强生成,并在推理阶段引入外部推理器进一步提升性能。最后,提出StructBench评测基准,包含超过1700个高难度实例,配套的StructScore评估指标采用多轮问答协议,精准衡量事实性准确度。对15个模型的评估表明,即使是领先的闭源系统仍表现不佳。我们的模型在编辑任务中表现突出,且推理阶段引入外部推理器在多种架构上均带来稳定提升。通过开源数据集、模型与评测工具,旨在推动结构化视觉内容的统一多模态基础发展。
原文摘要 · Abstract (English)
While modern visual generation models excel at creating aesthetically pleasing natural images, they struggle with producing or editing structured visuals like charts, diagrams, and mathematical figures, which demand composition planning, text rendering, and multimodal reasoning for factual fidelity. To address this, we present the first comprehensive, systematic investigation of this domain, encompassing data construction, model training, and an evaluation benchmark. First, we construct a large-scale dataset of 1.3 million high-quality structured image pairs derived from executable drawing programs and augmented with chain-of-thought reasoning annotations. Building on it, we train a unified model that integrates a VLM with FLUX.1 Kontext via a lightweight connector for enhanced multimodal understanding. A three-stage training curriculum enables progressive feature alignment, knowledge infusion, and reasoning-augmented generation, further boosted by an external reasoner at inference time. Finally, we introduce StructBench, a novel benchmark for generation and editing with over 1,700 challenging instances, and an accompanying evaluation metric, StructScore, which employs a multi-round Q\&A protocol to assess fine-grained factual accuracy. Evaluations of 15 models reveal that even leading closed-source systems remain far from satisfactory. Our model attains strong editing performance, and inference-time reasoning yields consistent gains across diverse architectures. By releasing the dataset, model, and benchmark, we aim to advance unified multimodal foundations for structured visuals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。