arXiv:2503.04050cs.CV2025-03被引 1

提升视觉任务中上下文学习能力的高效扩散模型

Underlying Semantic Diffusion for Effective and Efficient In-Context Learning

  • 分离并聚合适配器,增强多任务下语义理解与泛化
  • 反馈引导机制让模型更好捕捉边缘纹理等细节
  • 插件式采样策略使推理速度提升9.45倍

扩散模型在图像可控生成和密集预测任务中表现强劲,但现有方法难以有效捕捉底层语义(如边缘、纹理、形状)并利用上下文学习,限制了其上下文理解能力和生成质量。同时,高计算开销和慢推理速度制约其实时应用。为此,我们提出底层语义扩散模型(US-Diffusion),通过引入分离与聚合适配器(SGA),在共享架构下解耦不同任务输入条件,提升跨视觉域的上下文学习与泛化能力;提出反馈辅助学习(FAL)框架,利用反馈信号引导模型捕捉语义细节并动态适应任务特定上下文;设计即插即用的高效采样策略(ESS),在高噪声时间步优化密集采样,兼顾训练与推理效率。实验表明,US-Diffusion在Map2Image任务上平均降低7.47的FID,Image2Map任务上平均降低0.026的RMSE,推理速度提升约9.45倍,且在新数据集与任务中表现优异,展现强大鲁棒性与适应性。

原文摘要 · Abstract (English)

Diffusion models has emerged as a powerful framework for tasks like image controllable generation and dense prediction. However, existing models often struggle to capture underlying semantics (e.g., edges, textures, shapes) and effectively utilize in-context learning, limiting their contextual understanding and image generation quality. Additionally, high computational costs and slow inference speeds hinder their real-time applicability. To address these challenges, we propose Underlying Semantic Diffusion (US-Diffusion), an enhanced diffusion model that boosts underlying semantics learning, computational efficiency, and in-context learning capabilities on multi-task scenarios. We introduce Separate & Gather Adapter (SGA), which decouples input conditions for different tasks while sharing the architecture, enabling better in-context learning and generalization across diverse visual domains. We also present a Feedback-Aided Learning (FAL) framework, which leverages feedback signals to guide the model in capturing semantic details and dynamically adapting to task-specific contextual cues. Furthermore, we propose a plug-and-play Efficient Sampling Strategy (ESS) for dense sampling at time steps with high-noise levels, which aims at optimizing training and inference efficiency while maintaining strong in-context learning performance. Experimental results demonstrate that US-Diffusion outperforms the state-of-the-art method, achieving an average reduction of 7.47 in FID on Map2Image tasks and an average reduction of 0.026 in RMSE on Image2Map tasks, while achieving approximately 9.45 times faster inference speed. Our method also demonstrates superior training efficiency and in-context learning capabilities, excelling in new datasets and tasks, highlighting its robustness and adaptability across diverse visual domains.

扩散模型上下文学习高效推理多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。