利用多模态特征区域不确定性对齐,提升物体级异常检测能力
RUNA: Object-level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal Representations
- 通过双编码器捕捉上下文信息,设计区域不确定性对齐机制
- 在复杂场景下显著优于现有方法,尤其在多样物体实例中表现突出
- 适用于需要可靠异常识别的视觉系统,如自动驾驶
让目标检测器识别分布外(OOD)物体对构建可靠系统至关重要。主要挑战在于模型缺乏对陌生数据的监督信号,导致对OOD物体产生过度自信预测。尽管已有方法基于检测模型和分布内(ID)样本估计OOD不确定性,本文探索使用预训练视觉语言表征进行物体级OOD检测。首先分析了图像级CLIP-based方法在物体级场景中的局限性。在此基础上,提出RUNA框架:采用双编码器架构捕获丰富上下文信息,并引入区域不确定性对齐机制,有效区分ID与OOD物体。进一步提出少样本微调策略,对齐区域语义表征以增强相似物体的区分能力。实验表明,RUNA在物体级OOD检测上显著超越当前最优方法,尤其在包含多样化、复杂物体实例的挑战性场景中表现优异。
原文摘要 · Abstract (English)
Enabling object detectors to recognize out-of-distribution (OOD) objects is vital for building reliable systems. A primary obstacle stems from the fact that models frequently do not receive supervisory signals from unfamiliar data, leading to overly confident predictions regarding OOD objects. Despite previous progress that estimates OOD uncertainty based on the detection model and in-distribution (ID) samples, we explore using pre-trained vision-language representations for object-level OOD detection. We first discuss the limitations of applying image-level CLIP-based OOD detection methods to object-level scenarios. Building upon these insights, we propose RUNA, a novel framework that leverages a dual encoder architecture to capture rich contextual information and employs a regional uncertainty alignment mechanism to distinguish ID from OOD objects effectively. We introduce a few-shot fine-tuning approach that aligns region-level semantic representations to further improve the model's capability to discriminate between similar objects. Our experiments show that RUNA substantially surpasses state-of-the-art methods in object-level OOD detection, particularly in challenging scenarios with diverse and complex object instances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。