arXiv:2606.07638cs.CVcs.AI2026-06中稿 · the International …被引 1

让风景图生成更精准控制构图,支持摄影师级布局调节。

Anchor-Conditioned Compositional Control for Landscape Image Generation

论文配图:Anchor-Conditioned Compositional Control for Landscape Image Generation
图 1 · 摘自论文原文
  • 用四维构图向量+傅里叶编码注入扩散模型,实现精细布局控制。
  • 水平线检测率0.850,三分法对齐率达0.817,优于基线与消融实验。
  • 按场景类别分组训练可降低40%的水平线偏差,适合专业图像设计者。

尽管图像生成模型被广泛用作创作工具,但在构图控制方面仍远不如摄影师和视觉艺术家所掌握的精细程度。本文提出一种基于构图锚点的微调框架,用于风景图像生成。从训练图像中提取四维构图向量,并通过解耦交叉注意力机制结合傅里叶编码,以及三路无分类器引导丢弃,将其注入扩散模型。定量评估显示,该架构在基准模型及三种消融变体中达到最高的水平线检测率(0.850)和三分法对齐率(0.817)。类别特异性消融实验进一步表明,在构图同质的场景子集上训练,可比混合训练减少高达40%的水平线偏差,证明构图控制精度具有类别依赖性。

原文摘要 · Abstract (English)

Image generative models, though widely used as creative tools, offer limited support for the kind of compositional control that photographers and visual artists routinely exercise. This paper presents early results on an anchor conditioned finetuning framework for landscape image generation, in which a four dimensional compositional anchor vector is extracted from training images and injected into a diffusion model via a decoupled cross attention mechanism with Fourier encoding and three way classifier free guidance dropout. Quantitative evaluation against a baseline and three ablation variants shows that the proposed architecture achieves the highest horizon detection rate of 0.850 and the highest rule of thirds alignment of 0.817. A category specific ablation further demonstrates that training on compositionally homogeneous scene subsets reduces horizon deviation by up to 40 percent compared to mixed training. This establishes that compositional control precision is category dependent.

风景生成构图控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。