融合深度与视觉注意力,提升人眼注视目标检测精度
Leveraging Multi-Modal Saliency and Fusion for Gaze Target Detection
- 用单目深度估计构建3D图像表征,增强空间感知
- 多模态融合使在三个数据集上准确率超越现有方法
- 适合关注视觉注意力与人机交互的研究者
注视目标检测(GTD)旨在预测图像中人物的注视位置。该任务极具挑战性,需理解头部、身体与眼睛间关系及周围环境。本文提出一种新方法,通过融合图像中提取的多源信息实现检测。首先利用单目深度估计将2D图像映射为3D表示;随后生成融合深度信息的显著性模块图,突出被关注区域;同时提取人脸与深度模态特征,并最终融合所有模态以定位注视目标。我们在VideoAttentionTarget、GazeFollow和GOO-Real三个公开数据集上进行了定量评估,包含消融实验,结果表明该方法优于其他最先进方法,验证了其在注视目标检测中的有效性与潜力。
原文摘要 · Abstract (English)
Gaze target detection (GTD) is the task of predicting where a person in an image is looking. This is a challenging task, as it requires the ability to understand the relationship between the person's head, body, and eyes, as well as the surrounding environment. In this paper, we propose a novel method for GTD that fuses multiple pieces of information extracted from an image. First, we project the 2D image into a 3D representation using monocular depth estimation. We then extract a depth-infused saliency module map, which highlights the most salient (\textit{attention-grabbing}) regions in image for the subject in consideration. We also extract face and depth modalities from the image, and finally fuse all the extracted modalities to identify the gaze target. We quantitatively evaluated our method, including the ablation analysis on three publicly available datasets, namely VideoAttentionTarget, GazeFollow and GOO-Real, and showed that it outperforms other state-of-the-art methods. This suggests that our method is a promising new approach for GTD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。