通过思维链逐步演化风格,提升未知场景下的目标检测能力。
Style Evolving along Chain-of-Thought for Unknown-Domain Object Detection

- 用思维链逐步生成和融合多风格描述,模拟复杂环境
- 在五个恶劣天气场景中检测精度显著提升,优于基线方法
- 适合需要强泛化能力的自动驾驶、遥感等实际应用
近期提出的单域广义目标检测(Single-DGOD)任务旨在让检测器泛化到训练时未见过的多个未知领域。由于目标领域数据不可用,部分方法利用视觉-语言模型的多模态能力,通过文本提示估计跨域信息以增强泛化能力。这些方法通常采用单一文本提示,即一步提示法。然而,在处理雨夜等复杂风格组合时,其性能明显下降。原因在于许多场景涉及多种风格的叠加,一步提示难以有效融合多风格信息。为此,本文提出一种新方法——沿思维链演进风格(Style Evolving along Chain-of-Thought),通过逐步细化风格描述并引导风格多样化演化,实现风格信息的渐进式整合与扩展。该方法使模型能更准确地模拟各类风格特征,并逐步学习风格间的细微差异。同时,模型接触到分布不同的多样风格特征,增强了对未见领域的泛化能力。在五个恶劣天气场景及Real to Art基准测试中均取得显著性能提升。
原文摘要 · Abstract (English)
Recently, a task of Single-Domain Generalized Object Detection (Single-DGOD) is proposed, aiming to generalize a detector to multiple unknown domains never seen before during training. Due to the unavailability of target-domain data, some methods leverage the multimodal capabilities of vision-language models, using textual prompts to estimate cross-domain information, enhancing the model's generalization capability. These methods typically use a single textual prompt, often referred to as the one-step prompt method. However, when dealing with complex styles such as the combination of rain and night, we observe that the performance of the one-step prompt method tends to be relatively weak. The reason may be that many scenes incorporate not just a single style but a combination of multiple styles. The one-step prompt method may not effectively synthesize combined information involving various styles. To address this limitation, we propose a new method, i.e., Style Evolving along Chain-of-Thought, which aims to progressively integrate and expand style information along the chain of thought, enabling the continual evolution of styles. Specifically, by progressively refining style descriptions and guiding the diverse evolution of styles, this approach enables more accurate simulation of various style characteristics and helps the model gradually learn and adapt to subtle differences between styles. Additionally, it exposes the model to a broader range of style features with different data distributions, thereby enhancing its generalization capability in unseen domains. The significant performance gains over five adverse-weather scenarios and the Real to Art benchmark demonstrate the superiorities of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。