arXiv:2412.09050cs.CV2024-12被引 3

通过空间上下文增强,提升遮挡情况下人物交互识别准确率

ContextHOI: Spatial Context Learning for Human-Object Interaction Detection

  • 双分支结构分别提取物体特征与空间上下文信息
  • 在无额外标注下学习有效背景上下文,准确率超越现有方法
  • 特别适合处理遮挡或模糊实例的交互识别任务

空间上下文(如背景和周围环境)在人-物交互(HOI)识别中至关重要,尤其当前景对象模糊或被遮挡时。当前主流的HOI检测器基于检测变压器架构,虽能较好定位物体,但对空间上下文的利用不足,影响动作识别精度。为此,本文提出双分支框架ContextHOI,高效融合物体检测特征与空间上下文信息。在上下文分支中,模型无需人工标注背景即可学习有意义的空间上下文。此外,引入上下文感知的空间与语义监督,过滤无关噪声并捕捉关键上下文。在HICO-DET和v-coco基准上达到领先性能。为进一步验证,构建新基准HICO-ambiguous(HICO-DET中包含遮挡或受损实例的子集)。大量实验与可视化结果表明,ContextHOI在处理遮挡或模糊实例的交互识别方面显著提升性能。

原文摘要 · Abstract (English)

Spatial contexts, such as the backgrounds and surroundings, are considered critical in Human-Object Interaction (HOI) recognition, especially when the instance-centric foreground is blurred or occluded. Recent advancements in HOI detectors are usually built upon detection transformer pipelines. While such an object-detection-oriented paradigm shows promise in localizing objects, its exploration of spatial context is often insufficient for accurately recognizing human actions. To enhance the capabilities of object detectors for HOI detection, we present a dual-branch framework named ContextHOI, which efficiently captures both object detection features and spatial contexts. In the context branch, we train the model to extract informative spatial context without requiring additional hand-craft background labels. Furthermore, we introduce context-aware spatial and semantic supervision to the context branch to filter out irrelevant noise and capture informative contexts. ContextHOI achieves state-of-the-art performance on the HICO-DET and v-coco benchmarks. For further validation, we construct a novel benchmark, HICO-ambiguous, which is a subset of HICO-DET that contains images with occluded or impaired instance cues. Extensive experiments across all benchmarks, complemented by visualizations, underscore the enhancements provided by ContextHOI, especially in recognizing interactions involving occluded or blurred instances.

人物交互空间上下文遮挡识别检测框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。