重新评估物体可操作性分割的可复现性与尺度敏感性
Segmenting Object Affordances: Reproducibility and Sensitivity to Scale
- 构建可复现的实验框架,对比两种场景下的分割方法
- Mask2Former在多数测试集上表现最佳,但对尺度变化不鲁棒
- 提醒研究者注意训练与测试时物体分辨率差异的影响
视觉可操作性分割旨在识别物体上可交互的区域。现有方法多复用并改造语义分割的深度学习架构,且在小规模数据集上进行评估,但实验设置常不可复现,导致比较不公平且结果不一致。本文在两个单物体场景(无遮挡桌面物体和手持容器)下,采用可复现的实验设置对现有方法进行基准测试。我们重新训练了近期提出的Mask2Former模型用于可操作性分割,并发现其在两个场景的多数测试集上表现最优。分析表明,当物体分辨率与训练集不同时,模型性能显著下降,说明当前方法对尺度变化缺乏鲁棒性。
原文摘要 · Abstract (English)
Visual affordance segmentation identifies image regions of an object an agent can interact with. Existing methods re-use and adapt learning-based architectures for semantic segmentation to the affordance segmentation task and evaluate on small-size datasets. However, experimental setups are often not reproducible, thus leading to unfair and inconsistent comparisons. In this work, we benchmark these methods under a reproducible setup on two single objects scenarios, tabletop without occlusions and hand-held containers, to facilitate future comparisons. We include a version of a recent architecture, Mask2Former, re-trained for affordance segmentation and show that this model is the best-performing on most testing sets of both scenarios. Our analysis shows that models are not robust to scale variations when object resolutions differ from those in the training set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。