arXiv:2603.10061cs.RO2026-03被引 1

评估视觉语言模型在人机交互中的早期动作预测不确定性,提升机器人决策可靠性。

Decision-Aware Uncertainty Evaluation of Vision-Language Model-Based Early Action Anticipation for Human-Robot Interaction

  • 提出时序前缀评估框架,量化模型在部分观察下的置信度可靠性。
  • 发现模型在遮挡和视角变化下常过度自信,存在显著校准偏差。
  • 适用于需要可信置信度的机器人安全交互系统设计。

共享工作空间中的机器人需从部分、模糊的观测中理解人类动作,过度自信的早期预测可能导致不安全或干扰性交互。这一挑战在第一人称视角中尤为突出,因视角变化和遮挡增加了感知噪声与歧义。因此,下游人机交互模块不仅需要动作预测,还需在部分观测下获得可靠的置信度估计。近年来基于视觉语言模型的短时动作识别方法因其开放词汇和上下文感知推理能力被提出,但其在时序前缀阶段的不确定性可靠性尚未被系统研究。本文首次对基于视觉语言模型的短时动作识别在人机交互中的不确定性进行系统评估,提出时序前缀评估协议及校准性与选择性预测指标,并刻画了部分观测下的校准偏差模式与失败模式。研究为将视觉语言模型预测用于置信度门控的人机交互模块提供了缺失的可靠性证据。

原文摘要 · Abstract (English)

Robots in shared workspaces must interpret human actions from partial, ambiguous observations, where overconfident early predictions can lead to unsafe or disruptive interaction. This challenge is amplified in egocentric views, where viewpoint changes and occlusions increase perceptual noise and ambiguity. As a result, downstream human-robot interaction modules require not only an action hypothesis but also a trustworthy estimate of confidence under partial observation. Recent vision-language model-based approaches have been proposed for short-term action recognition due to their open-vocabulary and context-aware reasoning, but their uncertainty reliability in the temporal-prefix regime is largely uncharacterized. We present the first systematic evaluation of uncertainty in vision-language model-based short-term action recognition for human-robot interaction. We introduce a temporal-prefix evaluation protocol and metrics for calibration and selective prediction. We also characterize miscalibration patterns and failure modes under partial observations. Our study provides the missing reliability evidence needed to use vision-language model predictions in confidence-gated human-robot interaction modules.

人机交互动作预测不确定性评估视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。