基于人眼聚焦区域优化生成效率,实现高速高质图像视频生成
Foveated Diffusion: Efficient Spatially Adaptive Image and Video Generation
- 根据人眼视觉特性,对焦区高密度、周边低密度分配生成令牌
- 生成效果与全分辨率相当,令牌数减少超50%,速度显著提升
- 适合需要实时生成的交互式应用,如眼动追踪系统
扩散模型和流匹配模型在创意内容生成方面展现出前所未有的能力,例如交互式图像和流式视频生成。然而,对更高分辨率、帧率和上下文长度的需求使高效生成变得日益困难,因为计算复杂度随生成标记数呈二次增长。本文研究在用户注视位置已知或可估计(如通过眼动追踪)的场景下优化生成效率。利用人类视觉的离心率依赖性敏锐度:用户仅在注视点附近小区域内感知高分辨率信息,而视野外围的细节分辨能力迅速下降。提出一种建模视网膜分辨率的掩码,实现非均匀令牌分配——焦区高密度,周边低密度。在混合分辨率令牌设置下生成图像或视频,结果在感知上与全分辨率生成无异,同时大幅减少令牌数量和生成时间。为此,我们设计了一种从高分辨率数据直接构建混合分辨率令牌的原则性机制,使视网膜扩散模型可从现有基础模型后训练获得,且跨分辨率内容一致性得以保持。通过大量分析和精心设计的用户研究验证,证明视网膜机制是高效生成的实用且可扩展路径。
原文摘要 · Abstract (English)
Diffusion and flow matching models have unlocked unprecedented capabilities for creative content creation, such as interactive image and streaming video generation. The growing demand for higher resolutions, frame rates, and context lengths, however, makes efficient generation increasingly challenging, as computational complexity grows quadratically with the number of generated tokens. Our work seeks to optimize the efficiency of the generation process in settings where the user's gaze location is known or can be estimated, for example, by using eye tracking. In these settings, we leverage the eccentricity-dependent acuity of human vision: while a user perceives very high-resolution visual information in a small region around their gaze location (the foveal region), the ability to resolve detail quickly degrades in the periphery of the visual field. Our approach starts with a mask modeling the foveated resolution to allocate tokens non-uniformly, assigning higher token density to foveal regions and lower density to peripheral regions. An image or video is generated in a mixed-resolution token setting, yielding results perceptually indistinguishable from full-resolution generation, while drastically reducing the token count and generation time. To this end, we develop a principled mechanism for constructing mixed-resolution tokens directly from high-resolution data, allowing a foveated diffusion model to be post-trained from an existing base model while maintaining content consistency across resolutions. We validate our approach through extensive analysis and a carefully designed user study, demonstrating the efficacy of foveation as a practical and scalable axis for efficient generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。