arXiv:2603.06453cs.CV2026-03KDD

Pinterest自研图像生成系统,专攻电商场景的精准编辑与增强。

Pinterest Canvas: Large-Scale Image Generation at Pinterest

论文配图:Pinterest Canvas: Large-Scale Image Generation at Pinterest
图 1 · 摘自论文原文
  • 用多模态数据训练基础扩散模型,再针对性微调专用版本。
  • 背景增强和画幅扩展任务分别带来18.0%和12.5%的用户互动提升。
  • 支持多图合成、图生视频等复杂场景,适合高要求产品落地。

尽管近期图像生成模型在多样化任务上表现优异,但其灵活性导致仅靠提示词或简单推理调整难以控制,难以满足严格的产需要求。本文介绍Pinterest Canvas——一个为图像编辑与增强设计的大规模图像生成系统。该系统首先基于多样化的多模态数据集训练出具备广泛编辑能力的基础扩散模型;随后,不依赖单一通用模型,而是针对具体任务快速微调专用变体。我们阐述了Canvas的关键组件,并总结了数据集构建、训练与推理的最佳实践。通过背景增强与画幅外扩两个案例研究,展示了如何满足特定产品需求。线上A/B实验表明,优化后的图像分别带来18.0%和12.5%的显著互动提升;人工评估也验证了其优于第三方模型的表现。此外,我们还展示了多图场景合成与图像转视频等其他变体,证明该方法可泛化至多种下游任务。

原文摘要 · Abstract (English)

While recent image generation models demonstrate a remarkable ability to handle a wide variety of image generation tasks, this flexibility makes them hard to control via prompting or simple inference adaptation alone, rendering them unsuitable for use cases with strict product requirements. In this paper, we introduce Pinterest Canvas, our large-scale image generation system built to support image editing and enhancement use cases at Pinterest. Canvas is first trained on a diverse, multimodal dataset to produce a foundational diffusion model with broad image-editing capabilities. However, rather than relying on one generic model to handle every downstream task, we instead rapidly fine-tune variants of this base model on task-specific datasets, producing specialized models for individual use cases. We describe key components of Canvas and summarize our best practices for dataset curation, training, and inference. We also showcase task-specific variants through case studies on background enhancement and aspect-ratio outpainting, highlighting how we tackle their specific product requirements. Online A/B experiments demonstrate that our enhanced images receive a significant 18.0% and 12.5% engagement lift, respectively, and comparisons with human raters further validate that our models outperform third-party models on these tasks. Finally, we showcase other Canvas variants, including multi-image scene synthesis and image-to-video generation, demonstrating that our approach can generalize to a wide variety of potential downstream tasks.

图像生成扩散模型多模态工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。