arXiv:2605.20807cs.CV2026-05

通过预测结构图提升主体图像生成的细节保真度

Decomposing Subject-Driven Image Generation via Intermediate Structural Prediction

论文配图:Decomposing Subject-Driven Image Generation via Intermediate Structural Prediction
图 1 · 摘自论文原文
  • 先预测边缘图再渲染,分离结构与外观
  • 在100k图文对上训练,提升文本一致性
  • 适合需要高保真细节的图像生成任务

主体驱动的文本到图像生成仍难以保留高频率身份细节,如标志、图案和文字。现有方法通常直接在RGB空间操作,导致大幅编辑时细节退化。本文提出两阶段框架:先预测Canny边缘图,再基于源图像外观和预测结构生成最终图像。为改进文本处理,进一步构建了包含100,000对样本的全自动文本感知数据集,确保跨视角文本一致性。实验包括GPT-4.1评估和知识蒸馏研究,结果表明该方法优于多个基线,证明中间结构预测是实现高保真主体生成的有效路径。代码与数据集将公开。

原文摘要 · Abstract (English)

Subject-driven text-to-image generation still struggles to preserve high-frequency identity details such as logos, patterns, and text. Existing methods typically operate directly in RGB space, which often leads to detail degradation under substantial edits. We propose a two-stage framework that decouples structure from appearance by first predicting a Canny map and then rendering the final image conditioned on both the source appearance and the predicted structure. To improve text handling, we further introduce a fully automatic pipeline that constructs a 100k-pair text-aware dataset with cross-view textual consistency. Experiments, including GPT-4.1-based evaluation and a knowledge distillation study, show clear gains over selected baselines and suggest that intermediate structural prediction is an effective route for high-fidelity subject-driven generation. Our dataset and code will be made publicly available.

图像生成结构预测文本一致高保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。