arXiv:2604.27105cs.CV2026-04

用双流Transformer自动检测双摄像头下的互视与共同注意,提升心理学研究效率。

Automated Detection of Mutual Gaze and Joint Attention in Dual-Camera Settings via Dual-Stream Transformers

论文配图:Automated Detection of Mutual Gaze and Joint Attention in Dual-Camera Settings via Dual-Stream Transformers
图 1 · 摘自论文原文
  • 双流Transformer结合冻结的注视感知骨干网络,捕捉跨摄像头互动关系。
  • 在亲子互动数据集上表现优于卷积基线和主流多模态大模型。
  • 开源模型可适配不同实验室环境,助力行为科学自动化分析。

分析互视(MG)和共同注意(JA)对发展心理学至关重要,但传统依赖人工编码,耗时费力。在多摄像头实验环境下,自动化检测面临复杂的跨摄像头关系建模挑战。本文提出一种高效的双流Transformer架构,基于同步双摄像头视频,实现MG与JA的自动识别。方法采用冻结的注视感知骨干网络(GazeLLE)提取视觉先验,并引入定制化令牌融合机制,建模交互双方的空间与语义关系。在生态有效的照护者-婴儿互动数据集上评估,模型性能显著优于卷积基线和当前先进的多模态大语言模型(LLM)。通过开源模型及预训练权重,为行为科学家提供可微调的可扩展工具,有效弥合计算建模与应用交互研究之间的鸿沟。

原文摘要 · Abstract (English)

Analyzing mutual gaze (MG) and joint attention (JA) is critical in developmental psychology but traditionally relies on labor-intensive manual coding. Automating this process in multi-camera laboratory settings is computationally challenging due to complex cross-camera relational dynamics. In this paper, we propose a highly efficient dual-stream Transformer architecture for detecting MG and JA from synchronized dual-camera recordings. Our approach leverages frozen gaze-aware backbones (GazeLLE) to extract rich visual priors, combined with a custom token fusion mechanism to map the spatial and semantic relationships between interacting dyads. Evaluated on an ecologically valid dataset of caregiver-infant interactions, our model exhibits good performance, significantly outperforming both a convolutional baseline and a state-of-the-art multimodal Large Language Model (LLM). By open-sourcing our model and pre-trained weights, we provide behavioral scientists with a scalable tool that can be fine-tuned to diverse laboratory environments, effectively bridging the gap between computational modeling and applied interaction research.

互视检测双摄像头注意力分析行为建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。