arXiv:2409.01871cs.CVcs.AI2024-09被引 8

改进的卷积注意力模型,提升室内实时物体检测精度与速度。

Real-Time Indoor Object Detection based on hybrid CNN-Transformer Approach

  • 融合卷积与注意力机制,增强对复杂场景特征的识别能力。
  • 基于OpenImages v7构建32类室内物体专用数据集,提升检测针对性。
  • 兼顾精度与速度,适合增强现实等实时交互应用。

室内实时物体检测是计算机视觉中的挑战性领域,面临光照变化和背景复杂等难题。该研究评估现有数据集与模型,基于OpenImages v7构建专注于32个室内类别(如家具、电器)的新数据集,以提升真实应用场景的适配性。同时提出一种改进的CNN检测模型,引入注意力机制,强化对杂乱环境中关键特征的捕捉与优先级判断。实验表明,该方法在准确率与推理速度上均达到或超越现有先进模型水平,为实时室内物体检测开辟新路径。

原文摘要 · Abstract (English)

Real-time object detection in indoor settings is a challenging area of computer vision, faced with unique obstacles such as variable lighting and complex backgrounds. This field holds significant potential to revolutionize applications like augmented and mixed realities by enabling more seamless interactions between digital content and the physical world. However, the scarcity of research specifically fitted to the intricacies of indoor environments has highlighted a clear gap in the literature. To address this, our study delves into the evaluation of existing datasets and computational models, leading to the creation of a refined dataset. This new dataset is derived from OpenImages v7, focusing exclusively on 32 indoor categories selected for their relevance to real-world applications. Alongside this, we present an adaptation of a CNN detection model, incorporating an attention mechanism to enhance the model's ability to discern and prioritize critical features within cluttered indoor scenes. Our findings demonstrate that this approach is not just competitive with existing state-of-the-art models in accuracy and speed but also opens new avenues for research and application in the field of real-time indoor object detection.

室内检测CNN-Transformer实时检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。