首个面向人-物交互检测的鲁棒性评测基准,解决真实场景下模型性能下降问题。
RoHOI: Robustness Benchmark for Human-Object Interaction Detection
- 构建20类噪声类型评测集,覆盖真实世界多种干扰因素
- 现有模型在噪声下性能下降超40%,暴露严重脆弱性
- 提出语义感知渐进学习法,提升模型对部分/整体线索的适应能力
人-物交互(HOI)检测对人机协作至关重要,能实现情境感知支持。然而,现有模型在干净数据集上训练后,在真实环境中因未预见的噪声导致预测失准。为此,本文提出首个面向HOI检测的鲁棒性评测基准——RoHOI,系统评估模型在多样化挑战下的抗扰能力。该基准基于HICO-DET和V-COCO数据集,包含20种污染类型,并引入新的鲁棒性度量指标。我们对当前主流模型进行系统分析,发现其在各类噪声下性能显著下降。为提升鲁棒性,提出语义感知掩码渐进学习(SAMPL)策略,引导模型基于整体与局部线索动态优化,增强特征学习能力。大量实验表明,该方法优于现有最先进方法,树立了鲁棒HOI检测新标准。相关评测数据、数据集及代码已开源。
原文摘要 · Abstract (English)
Human-Object Interaction (HOI) detection is crucial for robot-human assistance, enabling context-aware support. However, models trained on clean datasets degrade in real-world conditions due to unforeseen corruptions, leading to inaccurate predictions. To address this, we introduce the first robustness benchmark for HOI detection, evaluating model resilience under diverse challenges. Despite advances, current models struggle with environmental variability, occlusions, and noise. Our benchmark, RoHOI, includes 20 corruption types based on the HICO-DET and V-COCO datasets and a new robustness-focused metric. We systematically analyze existing models in the HOI field, revealing significant performance drops under corruptions. To improve robustness, we propose a Semantic-Aware Masking-based Progressive Learning (SAMPL) strategy to guide the model to be optimized based on holistic and partial cues, thus dynamically adjusting the model's optimization to enhance robust feature learning. Extensive experiments show that our approach outperforms state-of-the-art methods, setting a new standard for robust HOI detection. Benchmarks, datasets, and code are available at https://github.com/KratosWen/RoHOI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。