arXiv:2608.27214cs.CV2026-08中稿 · ACM Multimedia 202…

解决多模态目标检测中未知物体误删问题,提升开放世界检测精度。

CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection

论文配图:CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection
图 1 · 摘自论文原文
  • 通过跨模态校准与动态抑制,优化已知类与未知类的判断边界。
  • 在Real-World Detection数据集上,U-mAP达21.7,K-mAP达40.8,超越此前最优结果。
  • 适合关注开放世界检测、多模态模型推理优化的研究者使用。

基于多模态基础模型的开放世界目标检测常因单向文本到视觉匹配导致语义模糊,且僵化的异常值惩罚会过度抑制靠近已知类别决策边界的未知物体。本文提出CODE(跨模态校准与动态抑制)框架,包含三个互补组件:跨模态联合置信度校准引入全局视觉原型,校准文本驱动的已知类别预测;不确定性引导的通用物体性增强通过局部视觉响应中的分类犹豫度,强化潜在未知物体;基于置信度差值的动态异常值抑制替代固定惩罚,保留分布外实例。在Real-World Detection基准测试中,采用OWL-ViT L/14作为主干网络,CODE在任务1中实现21.7 U-mAP和40.8 K-mAP,分别领先前序最佳方法2.6和2.3个百分点。

原文摘要 · Abstract (English)

Open World Object Detection (OWOD) built on multimodal foundation models often suffers from semantic ambiguity caused by unidirectional text-to-vision matching, while rigid outlier penalties may over-suppress unknown objects near known-class decision boundaries. We propose CODE (Cross-Modal Calibration and Dynamic Suppression), a unified inference-time framework with three complementary components. Cross-Modal Joint Confidence Calibration injects global visual prototypes to calibrate text-driven known-class predictions. Uncertainty-Guided Universal Objectness Enhancement measures classification hesitation from local visual responses to strengthen potential unknown objects. Dynamic Outlier Suppression via Confidence Margin replaces rigid suppression with a margin-aware adjustment that preserves ambiguous out-of-distribution instances. Experiments on the Real-World Detection benchmark demonstrate that, with the OWL-ViT L/14 backbone, CODE achieves 21.7 U-mAP and 40.8 K-mAP in Task 1, surpassing the previous state of the art by 2.6 and 2.3 points, respectively.

开放世界检测多模态目标检测动态抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。