arXiv:2502.06631cs.CVcs.AI2025-02中稿 · ICIP 2025 Workshop…被引 1

用置信度校准提升视觉语言模型的动作识别可靠性

Conformal Predictions for Human Action Recognition with Vision-Language Models

  • 引入置信度校准技术,减少动作识别候选类别数
  • 无需额外校准数据,通过调整softmax温度优化结果
  • 适合高风险场景下人机协作的可信AI系统

人类在环(HITL)系统在高风险真实应用中至关重要,需实现人工智能与人类决策者的协同。本文研究如何利用提供严格覆盖率保证的置信度校准(Conformal Prediction, CP)技术,增强基于视觉-语言模型(VLMs)的先进动作识别(HAR)系统的可靠性。实验表明,CP可显著减少平均候选类别数,且无需修改底层VLM。然而,此类方法常导致长尾分布,影响实际应用。为此,我们提出在不使用额外校准数据的前提下,通过调节softmax温度进行优化。该工作推动了多模态人机交互在动态真实环境中的发展。

原文摘要 · Abstract (English)

Human-in-the-Loop (HITL) systems are essential in high-stakes, real-world applications where AI must collaborate with human decision-makers. This work investigates how Conformal Prediction (CP) techniques, which provide rigorous coverage guarantees, can enhance the reliability of state-of-the-art human action recognition (HAR) systems built upon Vision-Language Models (VLMs). We demonstrate that CP can significantly reduce the average number of candidate classes without modifying the underlying VLM. However, these reductions often result in distributions with long tails which can hinder their practical utility. To mitigate this, we propose tuning the temperature of the softmax prediction, without using additional calibration data. This work contributes to ongoing efforts for multi-modal human-AI interaction in dynamic real-world environments.

动作识别置信度校准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。