arXiv:2503.14228cs.CVcs.AI2025-03

针对鱼眼图中人物小且旋转的问题,提出全景畸变感知的分块方法提升检测效果。

Panoramic Distortion-Aware Tokenization for Person Detection and Localization in Overhead Fisheye Images

  • 将鱼眼图重映射为全景图,利用其几何特性处理人物旋转与尺寸缩小问题。
  • 发现顶部人物在全景图中高度随垂直角度线性减小,据此设计分块策略。
  • 通过自相似分块和显著性保留机制,有效提升对小人物的检测能力。

俯视鱼眼图像中的人物检测因人物旋转和体型过小而具有挑战性。以往工作主要关注旋转问题,对小人物检测研究不足。本文将鱼眼图像重映射至等距柱状全景图,以处理旋转并利用全景几何特性更有效地应对小人物问题。传统检测方法倾向于关注较大目标,因其占据注意力图主导地位,导致小人物被忽略。在半球形等距柱状全景图中,我们发现人物在图像顶部附近,其视觉高度随垂直角度近似线性下降。基于此发现,提出全景畸变感知的分块策略:使用自相似图形划分全景特征,实现无间隙最优分割,并利用每块内最大显著性值来保留小人物的显著区域。提出一种结合全景重映射与分块策略的基于Transformer的人体检测与定位方法。大量实验表明,该方法在大规模数据集上优于现有常规方法。

原文摘要 · Abstract (English)

Person detection in overhead fisheye images is challenging due to person rotation and small persons. Prior work has mainly addressed person rotation, leaving the small-person problem underexplored. We remap fisheye images to equirectangular panoramas to handle rotation and exploit panoramic geometry to handle small persons more effectively. Conventional detection methods tend to favor larger persons because they dominate the attention maps, causing smaller persons to be missed. In hemispherical equirectangular panoramas, we find that apparent person height decreases approximately linearly with the vertical angle near the top of the image. Using this finding, we introduce panoramic distortion-aware tokenization to enhance the detection of small persons. This tokenization procedure divides panoramic features using self-similar figures that enable the determination of optimal divisions without gaps, and we leverage the maximum significance values in each tile of the token groups to preserve the significance areas of smaller persons. We propose a transformer-based person detection and localization method that combines panoramic-image remapping and the tokenization procedure. Extensive experiments demonstrated that our method outperforms conventional methods on large-scale datasets.

目标检测鱼眼图像全景图小目标

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。