融合全局场景信息提升汽车损伤检测精度
C-DiffDet+: Fusing Global Scene Context with Generative Denoising for High-Fidelity Car Damage Detection
- 用跨注意力机制融合全局场景与局部检测框特征
- 在CarDD数据集上超越现有模型,实现新基准
- 适合需要高精度细粒度检测的工业质检场景
在车辆损伤评估等挑战性视觉任务中,即使人类专家也难以可靠完成细粒度目标检测。尽管扩散模型DiffusionDet通过条件去噪实现了前沿性能,但在依赖上下文的场景中仍受限于局部特征的条件化。本文提出上下文感知融合(CAF)机制,利用跨注意力将全局场景上下文与局部候选框特征直接融合。全局上下文由独立专用编码器生成,捕捉全面环境信息,使每个目标候选框可关注场景级理解。该框架显著提升了生成式检测范式,实验表明在CarDD基准上优于现有模型,建立了细粒度领域中上下文感知目标检测的新性能标准。
原文摘要 · Abstract (English)
Fine-grained object detection in challenging visual domains, such as vehicle damage assessment, presents a formidable challenge even for human experts to resolve reliably. While DiffusionDet has advanced the state-of-the-art through conditional denoising diffusion, its performance remains limited by local feature conditioning in context-dependent scenarios. We address this fundamental limitation by introducing Context-Aware Fusion (CAF), which leverages cross-attention mechanisms to integrate global scene context with local proposal features directly. The global context is generated using a separate dedicated encoder that captures comprehensive environmental information, enabling each object proposal to attend to scene-level understanding. Our framework significantly enhances the generative detection paradigm by enabling each object proposal to attend to comprehensive environmental information. Experimental results demonstrate an improvement over state-of-the-art models on the CarDD benchmark, establishing new performance benchmarks for context-aware object detection in fine-grained domains
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。