arXiv:2503.20672cs.CV2025-03CVPR被引 17

构建超密集版式图文生成数据集,提升长文本商业内容生成质量

BizGen: Advancing Article-level Visual Text Rendering for Infographics Generation

  • 采用分层检索增强生成法构建65万张超密集版式图文数据集
  • 提出布局引导跨注意力机制,精准控制上百个子区域内容生成
  • 针对商业文案生成难题,适配幻灯片与信息图等实际场景

当前先进的文生图模型(如Flux、Ideogram 2.0)在句级图文渲染上已取得显著进展。本文聚焦更具挑战性的文章级图文渲染任务,提出基于用户提供的文章级描述和超密集版式,生成高质量商业内容(包括信息图与演示文稿)的新方法。核心挑战在于极长上下文长度以及高质量商业内容数据稀缺。与以往仅关注少数子区域和句级提示的工作不同,本研究需在包含数十至数百个子区域的复杂版式中保持精确布局对齐,难度显著提升。本文贡献两点:(i) 构建可扩展的高质量商业内容数据集Infographics-650K,通过分层检索增强生成方案实现,包含超密集版式与对应提示;(ii) 提出布局引导跨注意力机制,将多个区域级提示注入裁剪后的区域潜在空间,并在推理阶段利用布局条件化的CFG灵活优化每个子区域。实验表明,系统在自建的BizEval提示集上优于Flux与SD3等前沿模型,且通过充分消融实验验证各组件有效性。我们期望Infographics-650K与BizEval能推动社区在商业内容生成方向的发展。

原文摘要 · Abstract (English)

Recently, state-of-the-art text-to-image generation models, such as Flux and Ideogram 2.0, have made significant progress in sentence-level visual text rendering. In this paper, we focus on the more challenging scenarios of article-level visual text rendering and address a novel task of generating high-quality business content, including infographics and slides, based on user provided article-level descriptive prompts and ultra-dense layouts. The fundamental challenges are twofold: significantly longer context lengths and the scarcity of high-quality business content data. In contrast to most previous works that focus on a limited number of sub-regions and sentence-level prompts, ensuring precise adherence to ultra-dense layouts with tens or even hundreds of sub-regions in business content is far more challenging. We make two key technical contributions: (i) the construction of scalable, high-quality business content dataset, i.e., Infographics-650K, equipped with ultra-dense layouts and prompts by implementing a layer-wise retrieval-augmented infographic generation scheme; and (ii) a layout-guided cross attention scheme, which injects tens of region-wise prompts into a set of cropped region latent space according to the ultra-dense layouts, and refine each sub-regions flexibly during inference using a layout conditional CFG. We demonstrate the strong results of our system compared to previous SOTA systems such as Flux and SD3 on our BizEval prompt set. Additionally, we conduct thorough ablation experiments to verify the effectiveness of each component. We hope our constructed Infographics-650K and BizEval can encourage the broader community to advance the progress of business content generation.

图文生成信息图布局控制长文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。