用可控骨架生成数据精准的图像化图表,支持空间与主体双重控制。
ChArtist: Generating Pictorial Charts with Unified Spatial and Subject Control
- 基于骨架结构设计空间控制,保留数据信息同时灵活融合参考图像。
- 在3万组数据上训练,生成图表数据准确率显著提升。
- 适合需要精确数据可视化和创意表达的设计师与研究人员。
图像化图表是视觉叙事的有效媒介,能将视觉元素与数据图表自然结合。但其生成挑战在于视觉灵活性与图表结构刚性之间的冲突,需在保持数据准确性的同时兼顾视觉美感。现有方法依赖自然图像的密集结构线索(如边缘或深度图)作为条件信号,不适用于图像化图表生成。本文提出ChArtist,一种面向图像化图表生成的专用扩散模型,提供两种控制方式:1)与图表结构对齐的空间控制;2)尊重参考图像视觉特征的主题驱动控制。为此,我们引入基于骨架的空间控制表示,仅编码图表的数据信息,便于融入参考视觉且无刚性轮廓约束。基于扩散变压器(DiT)实现,并采用自适应位置编码管理双重控制。进一步提出空间门控注意力机制,调节空间与主题控制间的交互。为支持预训练模型微调,构建了包含3万组三元组(骨架、参考图像、图像化图表)的大规模数据集。还提出统一的数据准确度评估指标,衡量生成图表的数据忠实度。本工作表明,通过任务特异性表示,生成模型可实现数据驱动的视觉叙事。
原文摘要 · Abstract (English)
A pictorial chart is an effective medium for visual storytelling, seamlessly integrating visual elements with data charts. However, creating such images is challenging because the flexibility of visual elements often conflicts with the rigidity of chart structures. This process thus requires a creative deformation that maintains both data faithfulness and visual aesthetics. Current methods that extract dense structural cues from natural images (e.g., edge or depth maps) are ill-suited as conditioning signals for pictorial chart generation. We present ChArtist, a domain-specific diffusion model for generating pictorial charts automatically, offering two distinct types of control: 1) spatial control that aligns well with the chart structure, and 2) subject-driven control that respects the visual characteristics of a reference image. To achieve this, we introduce a skeleton-based spatial control representation. This representation encodes only the data-encoding information of the chart, allowing for the easy incorporation of reference visuals without a rigid outline constraint. We implement our method based on the Diffusion Transformer (DiT) and leverage an adaptive position encoding mechanism to manage these two controls. We further introduce Spatially Gated Attention to modulate the interaction between spatial control and subject control. To support the fine-tuning of pre-trained models for this task, we created a large-scale dataset of 30,000 triplets (skeleton, reference image, pictorial chart). We also propose a unified data accuracy metric to evaluate the data faithfulness of the generated charts. We believe this work demonstrates that current generative models can achieve data-driven visual storytelling by moving beyond general-purpose conditions to task-specific representations. Project page: https://chartist-ai.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。