arXiv:2507.10977cs.CVcs.AI2025-07中稿 · International Join…被引 1

提出新型多尺度小波注意力与射线编码,提升人物交互检测精度与效率。

Conceptualizing Multi-scale Wavelet Attention and Ray-based Encoding for Human-Object Interaction Detection

  • 设计小波注意力骨干网络,融合高低阶特征捕捉中间层次交互
  • 射线编码器通过可学习射线源强度优化注意力,减少计算开销
  • 在ImageNet与HICO-DET上实现更优性能,适合视觉理解任务研究者

人物交互(HOI)检测对于准确定位并描述人与物体之间的交互至关重要,有助于全面理解复杂视觉场景。然而,现有检测器常因依赖高资源消耗的训练方法和低效架构而难以实现可靠且高效的预测。为此,本文提出一种类小波注意力骨干网络与新颖的射线编码架构,专为HOI检测设计。小波骨干网络通过聚合来自不同卷积滤波器提取的低阶与高阶特征,弥补中间层次交互表达能力的不足。同时,射线编码器通过优化解码器对感兴趣区域的关注,实现多尺度注意力,并降低计算负担。借助可学习射线源的衰减强度,解码器能将查询嵌入与关键区域对齐,从而提升预测准确性。在ImageNet与HICO-DET等基准数据集上的实验结果验证了该架构的有效性。代码已公开于[https://github.com/henry-pay/RayEncoder]。

原文摘要 · Abstract (English)

Human-object interaction (HOI) detection is essential for accurately localizing and characterizing interactions between humans and objects, providing a comprehensive understanding of complex visual scenes across various domains. However, existing HOI detectors often struggle to deliver reliable predictions efficiently, relying on resource-intensive training methods and inefficient architectures. To address these challenges, we conceptualize a wavelet attention-like backbone and a novel ray-based encoder architecture tailored for HOI detection. Our wavelet backbone addresses the limitations of expressing middle-order interactions by aggregating discriminative features from the low- and high-order interactions extracted from diverse convolutional filters. Concurrently, the ray-based encoder facilitates multi-scale attention by optimizing the focus of the decoder on relevant regions of interest and mitigating computational overhead. As a result of harnessing the attenuated intensity of learnable ray origins, our decoder aligns query embeddings with emphasized regions of interest for accurate predictions. Experimental results on benchmark datasets, including ImageNet and HICO-DET, showcase the potential of our proposed architecture. The code is publicly available at [https://github.com/henry-pay/RayEncoder].

人机交互注意力机制多尺度建模目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。