用相关性检验发现:模型够强时,不确定性排序效果稳定,但难识别新场景。
Confidence-Gated Robot Autonomy: When Does Uncertainty Actually Help?

- 用斯皮尔曼相关和等效检验评估不确定性排序能力
- 模型能力不足时,不确定性排名弱且不稳定;能力强后三类方法表现相似
- 适合关注自主决策与容错的机器人系统研究者
机器人常利用预测不确定性决定是否自主执行或切换备用策略。在阈值门控自治中,不确定性主要通过排序错误的可能性来发挥作用。标准指标如期望校准误差和AUROC无法直接检验不确定性是否影响行动/延迟决策。因此,我们采用斯皮尔曼秩相关、配对自助等效检验及行动/延迟一致性评估。在三个时间活动识别基准上,发现存在一个数据集依赖的性能阈值,低于该阈值时不确定性提供的错误排序较弱且不稳定;高于该阈值时,软最大启发式、MC Dropout和集成方法产生相似的门控行为,而阈值选择对执行结果影响更大。多种子具身仿真显示,一旦匹配实际自主执行率,碰撞率和代价也呈现相同模式。在时间协变量偏移下,排序质量保持稳定,但细粒度语义分布外检测仍接近随机水平。结果表明,当基础模型足够胜任时,简单不确定性代理足以用于选择性门控,但无法用于语义新颖性检测。
原文摘要 · Abstract (English)
Robotic systems often use predictive uncertainty to decide whether to act autonomously or defer to a fallback policy. In threshold-gated autonomy, uncertainty matters mainly through its ability to rank likely errors. Standard metrics such as expected calibration error and AUROC do not directly test whether uncertainty changes act/defer decisions. We therefore evaluate uncertainty using Spearman rank correlation, paired bootstrap equivalence testing, and act/defer agreement. Across three temporal activity-recognition benchmarks, we find a dataset-dependent competence regime below which uncertainty provides a weak and unstable error ranking. Above this regime, softmax heuristics, MC Dropout, and ensembles produce similar gating behavior, while threshold choice has a much larger effect on execution outcomes. A multi-seed embodied simulation shows the same pattern for collision rate and cost once realized autonomy is matched. Under temporal covariate shift, ranking quality remains stable, but fine grained semantic OOD detection remains near chance. These results suggest that simple uncertainty proxies can suffice for selective gating once the base model is competent, but not for semantic novelty detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。