arXiv:2605.15908cs.CVcs.AI2026-05

提出可任意分辨率生成的像素扩散模型,突破传统网格限制。

RaPD: Resolution-Agnostic Pixel Diffusion via Semantics-Enriched Implicit Representations

论文配图:RaPD: Resolution-Agnostic Pixel Diffusion via Semantics-Enriched Implicit Representations
图 1 · 摘自论文原文
  • 在连续神经图像场潜空间中进行扩散生成,避免离散化瓶颈。
  • 单个去噪潜变量可任意分辨率渲染,且推理成本不变。
  • 适合需要高分辨率灵活生成的视觉生成任务。

自然图像具有连续性,但大多数生成模型在离散网格上合成图像,限制了分辨率灵活性。连续神经场虽能实现无分辨率渲染,但以往方法仅在解码阶段引入连续性作为插值模块,生成潜空间仍为离散且以重建为导向。本文提出RaPD(Resolution-agnostic Pixel Diffusion),在连续神经图像场(NIF)潜空间中执行扩散过程。通过语义表示引导实现生成感知的潜空间学习,以及坐标查询注意力渲染器实现坐标条件、尺度感知的渲染。单一去噪潜变量可通过改变查询坐标在任意分辨率下渲染,且扩散计算成本保持不变。实验表明,该方法在生成质量与分辨率可扩展性方面均表现优异。

原文摘要 · Abstract (English)

Natural images are continuous, yet most generative models synthesize them on discrete grids, limiting resolution-flexible generation. Continuous neural fields enable resolution-free rendering, but prior methods introduce continuity only at the decoding stage as an interpolation module, leaving the generative latent space discretized and reconstruction-oriented. We propose RaPD (Resolution-agnostic Pixel Diffusion), which performs diffusion in a continuous Neural Image Field (NIF) latent space. RaPD bridges this reconstruction-generation gap with Semantic Representation Guidance for generation-aware latent learning and a Coordinate-Queried Attention Renderer for coordinate-conditioned, scale-aware rendering. A single denoised latent can be rendered at arbitrary resolutions by changing only the query coordinates, keeping diffusion cost fixed. Experiments demonstrate superior generation quality and resolution scalability.

图像生成扩散模型连续表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。