arXiv:2603.23037cs.CVcs.AI2026-03

用可解释网络提升YOLOv10的可信度,让模型自知何时不可靠。

YOLOv10 with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection and trustworthy multimodal AI in computer vision perception

  • 用Kolmogorov-Arnold网络作为后验代理,分析检测置信度可靠性。
  • 在模糊、遮挡等场景下准确识别低可信预测,准确率提升显著。
  • 结合轻量级多模态接口,适合自动驾驶等高可靠需求场景。

本文提出一种基于Kolmogorov-Arnold网络的新框架,用于提升计算机视觉中目标检测的可解释性。该方法针对自动驾驶感知系统在视觉退化或模糊场景下置信度不可靠的问题,采用7个几何与语义特征,通过加法样条结构建模YOLOv10检测结果的可信度。该结构支持特征影响的直接可视化,生成平滑透明的功能映射,揭示模型信心是否充分。在COCO及巴斯大学校园图像数据集上的实验表明,该框架能准确识别模糊、遮挡或低纹理条件下的低可信预测,为过滤、审查或风险缓解提供可操作洞察。此外,通过自举语言-图像(BLIP)基础模型生成场景描述,构建轻量级多模态界面,不损害可解释层。最终系统实现可解释的目标检测与可信置信度估计,为自主与多模态人工智能应用提供透明实用的感知组件。

原文摘要 · Abstract (English)

The interpretable object detection capabilities of a novel Kolmogorov-Arnold network framework are examined here. The approach refers to a key limitation in computer vision for autonomous vehicles perception, and beyond. These systems offer limited transparency regarding the reliability of their confidence scores in visually degraded or ambiguous scenes. To address this limitation, a Kolmogorov-Arnold network is employed as an interpretable post-hoc surrogate to model the trustworthiness of the You Only Look Once (Yolov10) detections using seven geometric and semantic features. The additive spline-based structure of the Kolmogorov-Arnold network enables direct visualisation of each feature's influence. This produces smooth and transparent functional mappings that reveal when the model's confidence is well supported and when it is unreliable. Experiments on both Common Objects in Context (COCO), and images from the University of Bath campus demonstrate that the framework accurately identifies low-trust predictions under blur, occlusion, or low texture. This provides actionable insights for filtering, review, or downstream risk mitigation. Furthermore, a bootstrapped language-image (BLIP) foundation model generates descriptive captions of each scene. This tool enables a lightweight multimodal interface without affecting the interpretability layer. The resulting system delivers interpretable object detection with trustworthy confidence estimates. It offers a powerful tool for transparent and practical perception component for autonomous and multimodal artificial intelligence applications.

可解释性目标检测多模态AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。