通过动作与身份识别,从第三人称视频中定位第一人称拍摄者。
Sequence-Based Identification of First-Person Camera Wearers in Third-Person Views
- 基于时序特征融合运动与重识别信息进行定位
- 在同步双视角数据上实现92.3%的识别准确率
- 适合沉浸式学习与协作机器人场景研究
随着第一人称摄像头普及,共享环境中多摄像头交互研究日益重要。尽管Ego4D和Ego-Exo4D等大规模数据集推动了第一人称视觉发展,但多个摄像头佩戴者之间的交互仍研究不足,制约了沉浸式学习与协作机器人等应用。为此,我们提出了TF2025数据集,包含同步的第一人称与第三人称视角。同时,提出一种基于序列的方法,结合运动线索与人物重识别技术,实现对第三人称视频中第一人称佩戴者的精准定位。
原文摘要 · Abstract (English)
The increasing popularity of egocentric cameras has generated growing interest in studying multi-camera interactions in shared environments. Although large-scale datasets such as Ego4D and Ego-Exo4D have propelled egocentric vision research, interactions between multiple camera wearers remain underexplored-a key gap for applications like immersive learning and collaborative robotics. To bridge this, we present TF2025, an expanded dataset with synchronized first- and third-person views. In addition, we introduce a sequence-based method to identify first-person wearers in third-person footage, combining motion cues and person re-identification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。