arXiv:2509.22728cs.SDcs.AI2025-09被引 1

根据提示复杂度自动调节扩散模型生成参数,提升图像音频质量

Prompt-aware classifier free guidance for diffusion models

  • 基于提示语义和语言复杂度预测最优生成尺度
  • 在MSCOCO 2014和AudioCaps上显著提升保真度与对齐性
  • 无需训练即可优化预训练扩散模型,适合实际部署

扩散模型在图像和音频生成中取得显著进展,主要得益于无分类器引导(Classifier-Free Guidance)。然而,引导尺度的选择仍缺乏探索:固定尺度难以适应不同复杂度的提示,导致过饱和或对齐不足。为此,我们提出一种提示感知框架,可预测尺度依赖的质量并推理最优引导值。具体地,通过在多尺度下生成样本并用可靠评估指标打分,构建大规模合成数据集;一个轻量级预测器,以语义嵌入和语言复杂度为条件,估计多指标质量曲线,并通过带正则化的效用函数确定最佳尺度。在MSCOCO 2014和AudioCaps上的实验表明,该方法相比原始CFG持续提升保真度、对齐性和感知偏好。本工作证明,提示感知尺度选择是一种有效且无需训练的预训练扩散模型增强方案。

原文摘要 · Abstract (English)

Diffusion models have achieved remarkable progress in image and audio generation, largely due to Classifier-Free Guidance. However, the choice of guidance scale remains underexplored: a fixed scale often fails to generalize across prompts of varying complexity, leading to oversaturation or weak alignment. We address this gap by introducing a prompt-aware framework that predicts scale-dependent quality and selects the optimal guidance at inference. Specifically, we construct a large synthetic dataset by generating samples under multiple scales and scoring them with reliable evaluation metrics. A lightweight predictor, conditioned on semantic embeddings and linguistic complexity, estimates multi-metric quality curves and determines the best scale via a utility function with regularization. Experiments on MSCOCO~2014 and AudioCaps show consistent improvements over vanilla CFG, enhancing fidelity, alignment, and perceptual preference. This work demonstrates that prompt-aware scale selection provides an effective, training-free enhancement for pretrained diffusion backbones.

扩散模型生成质量提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。