让艺术化图表准确呈现数据与文字关系,避免生成错乱
ArtChart: Faithful Artistic Chart Generation with Integrated Text Rendering

- 先生成无文字灰度布局,再用控制模块保持数学结构
- 通过OCR和视觉语言模型奖励,提升文本准确性和版式正确性
- 适合需要精准数据可视化的设计师与研究者
艺术化图表将数据可视化与表现性笔触、纹理和排版结合,但对图像生成模型挑战巨大:输出仅在保持图表几何结构、精确的图像内文字以及标签与标记语义绑定时才可用。我们提出ArtChart框架,实现忠实的艺术化图表生成与集成文字渲染。给定结构化图表规范与艺术提示,ArtChart首先生成不带文字的灰度布局以编码目标图表几何,随后训练特定于图表的控制模块以保持数学结构。为解决剩余的文字与版式错误,进一步通过基于GRPO的强化学习优化生成策略,采用基于OCR的文本奖励、基于视觉语言模型的版式奖励及美学奖励。多专家蒸馏阶段将这些目标融合,提炼出统一的均衡生成策略。我们还构建了ArtChart-Bench——一个包含2000个双语提示的基准,覆盖四种图表类型、受控数值分布、多样标签/数值格式和15种艺术风格,并开发了六轴评估协议(数学逻辑、文本准确率、文本布局、美学、指令遵循、可读性)。在ArtChart-Bench上的实验表明,ArtChart始终优于仅用提示、图像编辑和通用ControlNet基线,在数学保真度和标签-布局绑定上提升显著,同时保持优异的视觉质量。结果表明,艺术化图表生成应被视为可靠的视觉传达,而非泛化的风格化图像合成。
原文摘要 · Abstract (English)
Artistic charts combine data visualization with expressive marks, textures, and typography, but they are difficult for image generators: an output is useful only when its stylization preserves chart geometry, exact in-image text, and the semantic binding between labels and marks. We introduce ArtChart, a framework for faithful artistic chart generation with integrated text rendering. Given a structured chart specification and an artistic prompt, ArtChart first renders a text-free grayscale layout that encodes the target chart geometry, then trains a chart-specific control module to preserve mathematical structure. To address the remaining text and layout errors, we further refine the generation policy through GRPO-based reinforcement learning with OCR-based text rewards, VLM-based layout rewards, and aesthetic rewards. A multi-expert distillation stage reconciles these objectives by distilling single-reward experts into one balanced generation policy. We also construct ArtChart-Bench, a bilingual 2K-prompt benchmark covering four chart types, controlled value distributions, diverse label/value formats, and 15 artistic styles, together with ArtChart-Eval, a six-axis evaluation protocol measuring mathematical logic, text accuracy, text layout, aesthetics, instruction following, and readability. Experiments on ArtChart-Bench show that ArtChart consistently outperforms prompt-only, image-editing, and generic ControlNet baselines, with the largest gains on mathematical fidelity and label-layout binding while maintaining competitive visual quality. These results suggest that artistic chart generation should be evaluated as reliable visual communication rather than as generic stylized image synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。