arXiv:2605.03614cs.CV2026-05被引 1

用贝叶斯视觉变压器提升物体可操作区域分割的准确性与可信度

Uncertainty Estimation in Instance Segmentation of Affordances via Bayesian Visual Transformers

论文配图:Uncertainty Estimation in Instance Segmentation of Affordances via Bayesian Visual Transformers
图 1 · 摘自论文原文
  • 采用样本和集成方法估算不确定性,扩展注意力架构解决新任务
  • 在IIT-Aff数据集上比确定性模型提升7.4个百分点的Fβ^w分数
  • 可同时分析语义与空间层面的不确定性,适合机器人交互等场景

视觉可操作性识别图像中具有潜在交互作用的区域,为场景理解提供新范式。准确且局部化的可操作性区域预测对自主机器人自然行动、人机交互、增强现实及假肢视觉设备至关重要。本文提出基于贝叶斯视觉变压器的实例分割方法,采用样本和集成策略进行不确定性估计。通过详细消融实验验证各组件影响,利用多子网络检测分布提取像素级的信念不确定性和随机不确定性。提出概率掩码质量度量,实现对概率实例分割模型中语义与空间变异的综合分析。结果表明,贝叶斯模型的全局共识显著提升掩码精细度与泛化能力,在挑战性IIT-Aff数据集上取得+7.4 p.p的Fβ^w分数提升。贝叶斯模型校准更优,减少过度置信概率,不确定性估计更可靠。定性结果显示,随机不确定性集中在物体轮廓,信念不确定性出现在视觉挑战像素,增强了模型可解释性。

原文摘要 · Abstract (English)

Visual affordances identify regions in an image with potential interactions, offering a novel paradigm for scene understanding. Recognizing affordances allows autonomous robots to act more naturally, could enhance human-robot interactions, enrich augmented reality systems, and benefit prosthetic vision devices. Accurate and localized prediction of affordance regions, rather than general saliency maps is crucial for these applications. We present a model for instance segmentation of affordances by adopting sample-based and ensembles approaches for uncertainty estimation. We extend an attention-based architecture for our novel task, showing with detailed ablation experiments the effects of each component. By comparing the distribution of these different detections, we extract pixel-wise epistemic and aleatoric variances at both the semantic and spatial levels. In addition, we propose a novel measure called Probability-based Mask Quality, which enables a comprehensive analysis of semantic and spatial variations in a probabilistic instance segmentation model. Our results show that the global consensus of multiple sub-networks of Bayesian models improve deterministic networks due to a better mask refinement and generalization. This fact, joined with the more powerful features extracted by attention-based mechanisms, represent an improvement of +7.4 p.p on the $F_β^w$ score in the challenging IIT-Aff dataset. Bayesian models are also better calibrated, producing less overconfident probabilities and with a better uncertainty estimation. Qualitative results show that aleatoric variance appears in the contour of the objects, while the epistemic variance is observed in visual challenging pixels, adding interpretability to the neural network.

实例分割不确定性估计贝叶斯模型视觉变压器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。