提升文本引导计数在恶劣图像下的鲁棒性,无需标注也能自适应修复退化图像。
Test-Time Training for Robust Text-Guided Open-Vocabulary Object Counting

- 测试时训练双架构,仅更新轻量去噪模块,保持原计数结构不变。
- 在六类退化条件下(雨、雾、暗光等)显著降低计数误差,最高提升17.3%。
- 适合部署在真实场景中需快速适应恶劣视觉条件的开放词汇计数系统。
文本引导的开放词汇物体计数(TOOC)可对任意文本提示指定的物体类别进行计数,相比传统封闭集计数更具灵活性。然而,现有方法主要在理想图像上开发与评估,而真实场景常受雨、雾、暗光、高斯噪声、椒盐噪声及混合退化影响,严重降低视觉质量并破坏视觉-语言对齐。为此,我们提出首个针对退化条件下的鲁棒性评估基准Robust-TOOC,涵盖六种典型退化类型。为提升鲁棒性同时保留原有计数架构,我们设计了Dual-TTT——一种双架构测试时训练框架。测试阶段,Dual-TTT仅更新文本引导轻量去噪模块(TL-Denoiser),冻结原始计数网络。受扩散模型启发,该模块优化以去除退化条件下的特征噪声。由于仅在测试时微调去噪模块,Dual-TTT无需额外标注,可无缝集成至现有TOOC模型而不修改其架构。在多个近期TOOC基线上的大量实验验证了方法的有效性。
原文摘要 · Abstract (English)
Text-guided Open-vocabulary Object Counting (TOOC) enables counting arbitrary object categories specified by text prompts, offering substantially greater flexibility than conventional closed-set counting. However, existing TOOC methods are developed and evaluated primarily on ideal images, while real-world scenes often suffer from adverse conditions such as rain, fog, darkness, and sensor noise, which severely degrade visual quality and impair vision-language alignment. To bridge this gap, we introduce Robust-TOOC, the first benchmark for evaluating TOOC under diverse corruption conditions, which covers six representative degradation types: rain, fog, darkness, Gaussian noise, salt-and-pepper noise, and mixed corruption. To improve robustness while preserving the original counting architecture, we propose Dual-TTT, a dual-architecture test-time training framework for TOOC. Specifically, during test-time training, Dual-TTT updates only the Text-guided Lightweight Denoising module (TL-Denoiser), while keeping the original counting network frozen. Inspired by diffusion models, the TL-Denoiser is optimized to remove corruption-aware noise from image representations under degraded conditions. Since only the TL-Denoiser is trained at test time, Dual-TTT is annotation-free and can be seamlessly integrated into existing TOOC models without modifying their original architecture. Extensive experiments on multiple recent TOOC baselines demonstrate the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。