首个HOI检测鲁棒性评估基准,提升模型在分布偏移下的表现。
On the Robustness of Human-Object Interaction Detection against Distribution Shift
- 构建首个自动化生成的HOI鲁棒性评测基准
- 40+模型测试显示现有方法在分布偏移下表现不足
- 提出可插拔的数据增强与特征融合策略,通用性强
近年来,人-物体交互(HOI)检测取得了显著进展,但现有工作多聚焦于理想图像与自然分布的标准设置,难以应对实际场景中不可避免的分布偏移问题,限制了其应用。本文通过构建、分析并增强HOI检测模型在多种分布偏移下的鲁棒性,首次提出一种自动化的基准构建方法,对超过40种现有HOI检测模型进行了评估,揭示了它们在分布偏移下的不足,并分析了不同框架的特点及HOI任务鲁棒性的特殊性。基于这些洞察,本文提出两种简单、可插拔的改进方法:(1)结合mixup的跨域数据增强,(2)使用冻结视觉基础模型的特征融合策略。实验表明,该方法显著提升了多种模型的鲁棒性,且在标准基准上也带来增益。相关数据集与代码将公开。
原文摘要 · Abstract (English)
Human-Object Interaction (HOI) detection has seen substantial advances in recent years. However, existing works focus on the standard setting with ideal images and natural distribution, far from practical scenarios with inevitable distribution shifts. This hampers the practical applicability of HOI detection. In this work, we investigate this issue by benchmarking, analyzing, and enhancing the robustness of HOI detection models under various distribution shifts. We start by proposing a novel automated approach to create the first robustness evaluation benchmark for HOI detection. Subsequently, we evaluate more than 40 existing HOI detection models on this benchmark, showing their insufficiency, analyzing the features of different frameworks, and discussing how the robustness in HOI is different from other tasks. With the insights from such analyses, we propose to improve the robustness of HOI detection methods through: (1) a cross-domain data augmentation integrated with mixup, and (2) a feature fusion strategy with frozen vision foundation models. Both are simple, plug-and-play, and applicable to various methods. Our experimental results demonstrate that the proposed approach significantly increases the robustness of various methods, with benefits on standard benchmarks, too. The dataset and code will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。