剖析两阶段人物交互检测模型的失败模式,揭示复杂场景下的推理缺陷。
A Study of Failure Modes in Two-Stage Human-Object Interaction Detection

- 拆解交互场景维度,从多人、共用物体等配置分析模型表现
- 发现高整体准确率下仍存在严重推理错误,尤其在罕见组合中
- 适合关注视觉推理局限与模型鲁棒性的研究者参考
人-物交互(HOI)检测旨在识别图像中人物与物体间的交互行为。尽管近期进展提升了基准测试的性能,但现有评估多集中于整体预测准确率,对模型失败原因缺乏深入理解。尤其在涉及多人和稀有交互组合的复杂场景中,现代模型常表现不佳。本文针对两阶段HOI模型展开研究,不构建大规模新基准,而是将其分解为多个可解释的视角,通过分析不同人-物-交互配置下的模型行为,识别各类失败模式。我们从现有HOI数据集中选取子集,按多人交互、物体共享等配置组织图像,系统考察模型在不同场景构成下的表现及失败原因。结果表明,高整体性能并不意味着对人-物关系具备稳健的视觉推理能力。本研究为理解HOI模型局限性提供了重要洞见,有助于未来研究方向的优化。
原文摘要 · Abstract (English)
Human-object interaction (HOI) detection aims to detect interactions between humans and objects in images. While recent advances have improved performance on existing benchmarks, their evaluations mainly focus on overall prediction accuracy and provide limited insight into the underlying causes of model failures. In particular, modern models often struggle in complex scenes involving multiple people and rare interaction combinations. In this work, we present a study to better understand the failure modes of two-stage HOI models, which form the basis of many current HOI detection approaches. Rather than constructing a large-scale benchmark, we instead decompose HOI detection into multiple interpretable perspectives and analyze model behavior across these dimensions to study different types of failure patterns. We curate a subset of images from an existing HOI dataset organized by human-object-interaction configurations (e.g., multi-person interactions and object sharing), and analyze model behavior under these configurations to examine different failure modes. This design allows us to analyze how these HOI models behave under different scene compositions and why their predictions fail. Importantly, high overall benchmark performance does not necessarily reflect robust visual reasoning about human-object relationships. We hope that this study can provide useful insights into the limitations of HOI models and offer observations for future research in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。