多步梯度优化提升边缘设备图像修复质量
Multi-Step Guided Diffusion for Image Restoration on Edge Devices: Toward Lightweight Perception in Embodied AI
- 每步去噪中多次梯度更新,增强恢复精度
- 超分辨率与去模糊任务中PSNR和LPIPS显著提升
- 适配无人机等移动机器人,实时感知更轻量
扩散模型在无需特定任务重训练的情况下,表现出解决逆问题的出色灵活性。然而,现有方法如流形保持引导扩散(MPGD)仅在每步去噪中进行一次梯度更新,限制了修复质量与鲁棒性,尤其在嵌入式或分布外场景下。本文提出在每个去噪时间步内引入多步优化策略,显著提升图像质量、感知准确性与泛化能力。在超分辨率与高斯模糊去除任务上的实验表明,每步增加梯度更新次数可有效提升PSNR与LPIPS,且延迟开销极小。我们在Jetson Orin Nano上验证了该方法,使用退化的ImageNet与无人机数据集,证明原基于人脸数据训练的MPGD能有效泛化至自然与航拍场景。结果表明,MPGD具备作为轻量级即插即用修复模块的潜力,适用于无人机、移动机器人等具身智能体的实时视觉感知。
原文摘要 · Abstract (English)
Diffusion models have shown remarkable flexibility for solving inverse problems without task-specific retraining. However, existing approaches such as Manifold Preserving Guided Diffusion (MPGD) apply only a single gradient update per denoising step, limiting restoration fidelity and robustness, especially in embedded or out-of-distribution settings. In this work, we introduce a multistep optimization strategy within each denoising timestep, significantly enhancing image quality, perceptual accuracy, and generalization. Our experiments on super-resolution and Gaussian deblurring demonstrate that increasing the number of gradient updates per step improves LPIPS and PSNR with minimal latency overhead. Notably, we validate this approach on a Jetson Orin Nano using degraded ImageNet and a UAV dataset, showing that MPGD, originally trained on face datasets, generalizes effectively to natural and aerial scenes. Our findings highlight MPGD's potential as a lightweight, plug-and-play restoration module for real-time visual perception in embodied AI agents such as drones and mobile robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。