分析视觉退化下操作理解中关键关系的可靠性,发现抓取等动态关系最脆弱。
Trustworthy Visual Predicates for Robust Manipulation Understanding under Degradation

- 构建关系谓词可靠性框架,量化退化对接触、抓取等关系的影响
- 严重退化下谓词错误使理解准确率从0.89降至0.58
- 适合关注机器人感知鲁棒性与视觉推理诊断的研究者
操作理解依赖可靠的视觉谓词,如接触、支撑、包含、运动耦合、抓取、释放及主动手参与。尽管这些谓词广泛用于事件链、图模型和神经符号系统,但其在视觉退化下的可靠性鲜有直接分析。本文提出谓词级可靠性框架,应对模糊、遮挡、光照变化、低分辨率、帧丢失和检测噪声等问题。该框架定义了结构化谓词词汇表,实现置信度感知的谓词估计,并引入五类可靠性度量:谓词保持性、退化敏感性、时间一致性、置信度加权稳定性及下游影响。在受控操作视频及公开的自指或双人数据集(包括VISOR/EPIC-KITCHENS、H2O、ARCTIC)上的实验表明,谓词失效具有结构性而非随机性:静态空间谓词相对稳健,而接触敏感、动态及衍生谓词(如抓取、释放)更易受损。严重退化下,检测噪声、遮挡和帧丢失导致最强可靠性损失。下游分析显示,退化谓词使操作理解准确率从0.89降至0.58;去除置信度加权在中度退化下使准确率从0.74降至0.64。结果表明,谓词可靠性可作为视觉感知与结构化操作推理之间的诊断层。
原文摘要 · Abstract (English)
Manipulation understanding requires reliable relational evidence, such as contact, support, containment, motion coupling, grasp, release, and active-hand involvement. Although these visual predicates are widely used in event-chain, graph-based, and neuro-symbolic models, their reliability under visual degradation is rarely analyzed directly. This paper introduces a predicate-level reliability framework for robust manipulation understanding under blur, occlusion, illumination change, low resolution, frame dropping, and detection noise. The framework defines a structured predicate vocabulary, confidence-aware predicate estimation, and reliability metrics for predicate preservation, degradation sensitivity, temporal consistency, confidence-weighted stability, and downstream impact. Experiments on controlled manipulation videos and public egocentric or bimanual datasets, including VISOR/EPIC-KITCHENS, H2O, and ARCTIC, show that predicate failures are structured rather than uniform. Static spatial predicates remain comparatively robust, whereas contact-sensitive, dynamic, and derived predicates such as grasp and release are more fragile. Under severe degradation, detection noise, occlusion, and frame dropping cause the strongest reliability losses. Downstream analysis shows that degraded predicates reduce manipulation-understanding accuracy from 0.89 to 0.58, while removing confidence weighting under moderate degradation reduces accuracy from 0.74 to 0.64. These results show that predicate reliability provides a diagnostic layer between visual perception and structured manipulation reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。