arXiv:2603.19676cs.CVcs.AI2026-03

让扩散模型准确生成指定数量物体,无需重训练。

ATHENA: Adaptive Test-Time Steering for Improving Count Fidelity in Diffusion Models

  • 采样过程中动态估计物体数量并修正噪声,提前引导生成方向。
  • 在高数量目标场景下,计数准确率显著提升,最高达85%以上。
  • 无需修改模型或重新训练,适配多种扩散模型使用。

文本到图像的扩散模型虽具备高视觉保真度,但在提示中明确指定物体数量时却存在系统性失败。为此,我们提出ATHENA——一种模型无关、测试时自适应的调控框架,可在不修改模型结构或重训练的前提下提升物体计数保真度。ATHENA利用采样过程中的中间表示估算物体数量,并在去噪早期施加计数感知的噪声修正,使生成轨迹在结构错误固化前得到引导。我们设计了三种逐步增强的ATHENA变体,从基于提示的静态调控到动态调整的计数感知控制,计算开销与精度呈可调节权衡。在多个基准数据集和新构建的视觉语义复杂数据集上的实验表明,ATHENA能持续提升计数保真度,尤其在较高目标数量下表现突出,且在多种扩散主干网络上保持良好的精度-速度平衡。

原文摘要 · Abstract (English)

Text-to-image diffusion models achieve high visual fidelity but surprisingly exhibit systematic failures in numerical control when prompts specify explicit object counts. To address this limitation, we introduce ATHENA, a model-agnostic, test-time adaptive steering framework that improves object count fidelity without modifying model architectures or requiring retraining. ATHENA leverages intermediate representations during sampling to estimate object counts and applies count-aware noise corrections early in the denoising process, steering the generation trajectory before structural errors become difficult to revise. We present three progressively more advanced variants of ATHENA that trade additional computation for improved numerical accuracy, ranging from static prompt-based steering to dynamically adjusted count-aware control. Experiments on established benchmarks and a new visually and semantically complex dataset show that ATHENA consistently improves count fidelity, particularly at higher target counts, while maintaining favorable accuracy-runtime trade-offs across multiple diffusion backbones.

扩散模型计数生成测试时调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。