arXiv:2606.17342cs.CV2026-06

用扩散模型学习纹理的熵最大概率模型,生成效果更优且可平滑插值。

Learning a Maximum Entropy Model for Visual Textures using Diffusion

论文配图:Learning a Maximum Entropy Model for Visual Textures using Diffusion
图 1 · 摘自论文原文
  • 基于扩散模型无监督学习纹理统计量,构建熵最大概率模型。
  • 仅用512个统计量生成质量媲美甚至超越177k统计量的现役模型。
  • 模型表征空间轨迹可生成平滑过渡的连续纹理,适合可控生成任务。

视觉纹理——包含重复元素的空间均匀图像区域(如草地、树皮)——在视觉场景中普遍存在,对材料和物体识别具有重要意义。现有纹理模型从单张纹理图像中提取关键统计量,并通过匹配这些统计量生成高质量样本。然而,其统计量要么人为设计,要么基于为其他任务(如物体识别)预训练的网络。本文提出首个无监督学习统计量的严谨方法,用于约束最大熵概率模型。我们借助生成扩散模型的方法推导出训练与采样流程,并与传统统计匹配方法对比。尽管模型紧凑(仅512个统计量),生成的纹理质量与当前最优模型(约17.7万统计量)相当或更优。通过合成在一种模型下难以区分但在另一种模型下差异最大的图像,揭示了两者的相对优劣。此外,不同于以往统计纹理模型,本模型在表示空间中的直线路径能生成平滑插值的同质纹理样本。

原文摘要 · Abstract (English)

Visual textures -- spatially homogeneous image regions containing repeated elements (e.g. a field of grass, the bark of a tree) -- are ubiquitous in visual scenes and provide important cues for recognizing and analyzing materials and objects. A number of existing texture models extract essential statistics from a single texture image, and can then generate high-quality samples that are visually similar to the original by matching these statistics. However, their statistics are either hand-designed or based on a network pretrained for another purpose (e.g., object recognition). Here, we develop the first principled method for unsupervised learning of a set of statistics that are used to constrain a maximum entropy probability model. We leverage methods developed for generative diffusion models to derive training and sampling procedures, and compare these to the traditional method of sampling via matching the statistics. Despite the compactness of our trained model (512 statistics), it generates texture images whose quality is as good as or better than the current state-of-the-art model (~177k statistics). A more direct comparison of the two models, obtained by synthesizing images that are indistinguishable for one model but maximally different for the other, reveals their relative strengths and weaknesses. Finally, we show that unlike previous statistical texture models, a straight trajectory in the representation space of our model generates homogeneous texture samples that interpolate smoothly between the features of the two end points.

纹理生成扩散模型最大熵无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。