让图表既好看又准确,自动合成艺术化数据图
Semantic-Structural Alignment for Generative Pictorial Charts

- 用文本和结构图双重控制生成过程
- 在多种视觉通道上保持数据结构一致
- 适合需要美观数据可视化的研究者
传统统计图表精确但缺乏视觉吸引力。本文提出一种生成式框架,自动合成兼具语义表达与结构忠实性的象形图表。不同于简单图像风格化,该方法将问题建模为双条件生成任务:文本提示捕捉编辑意图的语义上下文,上下文图像提供抽象统计图表的全局结构。在多模态扩散变换器中,引入两个互补的特征级机制:结构对齐以锚定空间布局,语义对齐以传递参考图像的表达性纹理。该方法适用于长度、面积、角度和位置等多种视觉通道及多样语义领域,生成的艺术性强且结构一致的图表。大量定量评估与感知用户研究表明,本框架优于传统可控生成与图像编辑基线,为表达性视觉叙事中的高保真数据驱动生成建模奠定基础。
原文摘要 · Abstract (English)
Traditional statistical graphics are precise but often lack the visual appeal, memorability, and engagement of pictorial charts. We present a generative framework for the automated synthesis of pictorial charts that bridges the gap between semantic expression and structural faithfulness. Rather than treating charts merely as images to be stylized, we frame the problem as a dual-conditioned generation task guided by two parallel external control signals: a text prompt capturing the semantic context of the editing intent, and a context image providing the abstract statistical chart's global structure. To reinforce these controls within a Multi-Modal Diffusion Transformer, we introduce two complementary feature-level mechanisms: structural alignment to anchor spatial layouts to the input chart, and semantic alignment to transfer expressive textures from reference images. Generalizing across major visual channels (i.e., length, area, angle, and position) and diverse semantic domains, our method produces pictorial charts that are both artistically compelling and structurally consistent. Extensive quantitative evaluations and perceptual user studies demonstrate that our framework outperforms traditional controllable generation and image editing baselines, providing a foundation for high-fidelity, data-driven generative modeling in expressive visual storytelling. Project page: https://ssalign.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。