arXiv:2510.06295cs.CVcs.AI2025-10被引 3

MobilePicasso实现4K图像高效编辑,低内存高画质。

Efficient High-Resolution Image Editing with Hallucination-Aware Loss and Adaptive Tiling

  • 用抗幻觉损失在标准分辨率编辑图像
  • 通过自适应分块上采样提升画质,降低51%幻觉
  • 适合移动端部署,比云端模型还快

高分辨率(4K)图像到图像生成在移动应用中日益重要。现有扩散模型在资源受限设备上进行图像编辑时面临内存和画质的显著挑战。本文提出MobilePicasso,一种新系统,可在保持低计算成本和内存使用的情况下实现高效高分辨率图像编辑。MobilePicasso包含三个阶段:(i) 在标准分辨率下使用抗幻觉损失进行图像编辑;(ii) 通过潜在空间投影避免像素空间操作;(iii) 利用自适应上下文保持分块方法将编辑后的潜在表示上采样至更高分辨率。46名用户的实验表明,MobilePicasso在图像质量上提升18-48%,幻觉减少14-51%。其延迟显著降低,例如最高达55.8倍加速,运行内存仅增加9%。令人惊讶的是,MobilePicasso在设备端的运行速度甚至超过基于A100 GPU的服务器端高分辨率图像编辑模型。

原文摘要 · Abstract (English)

High-resolution (4K) image-to-image synthesis has become increasingly important for mobile applications. Existing diffusion models for image editing face significant challenges, in terms of memory and image quality, when deployed on resource-constrained devices. In this paper, we present MobilePicasso, a novel system that enables efficient image editing at high resolutions, while minimising computational cost and memory usage. MobilePicasso comprises three stages: (i) performing image editing at a standard resolution with hallucination-aware loss, (ii) applying latent projection to overcome going to the pixel space, and (iii) upscaling the edited image latent to a higher resolution with adaptive context-preserving tiling. Our user study with 46 participants reveals that MobilePicasso not only improves image quality by 18-48% but reduces hallucinations by 14-51% over existing methods. MobilePicasso demonstrates significantly lower latency, e.g., up to 55.8$\times$ speed-up, yet with a small increase in runtime memory, e.g., a mere 9% increase over prior work. Surprisingly, the on-device runtime of MobilePicasso is observed to be faster than a server-based high-resolution image editing model running on an A100 GPU.

图像编辑扩散模型移动端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。