解决虚拟会议中参会者跨画面跳跃跟踪难的问题。
Tracking Virtual Meetings in the Wild: Re-identification in Multi-Participant Virtual Meetings

- 利用时空先验建模虚拟会议窗口移动规律
- 相比传统方法错误率降低95%
- 适合研究远程会议分析与身份追踪的场景
近年来,工作场所和教育机构广泛采用虚拟会议平台,催生了对会议内容分析与洞察提取的迫切需求,这要求有效检测并追踪个体。然而,视频会议的录制布局缺乏标准化,且不同平台采集方式各异,导致数据流获取与统一分析面临挑战。本文提出一种针对最普遍视频录制形式——单源视频流中多参与者网格布局(如图1所示)——的解决方案,无需依赖参与者位置元数据,且对数据采集方式假设极少。传统方法常结合YOLO模型与追踪算法,假设线性运动轨迹,类似闭路电视监控场景,但在虚拟会议中,参与者窗口可能突然切换位置,频繁进出会议导致非线性运动,破坏依赖连续运动的光流追踪方法,造成多个参与者被误分配至同一追踪器。本文提出一种新方法,通过挖掘领域内特有的时空先验信息来提升追踪能力。实验表明,该方法相较基于YOLO的追踪基线平均错误率降低95%。
原文摘要 · Abstract (English)
In recent years, workplaces and educational institutes have widely adopted virtual meeting platforms. This has led to a growing interest in analyzing and extracting insights from these meetings, which requires effective detection and tracking of unique individuals. In practice, there is no standardization in video meetings recording layout, and how they are captured across the different platforms and services. This, in turn, creates a challenge in acquiring this data stream and analyzing it in a uniform fashion. Our approach provides a solution to the most general form of video recording, usually consisting of a grid of participants (\cref{fig:videomeeting}) from a single video source with no metadata on participant locations, while using the least amount of constraints and assumptions as to how the data was acquired. Conventional approaches often use YOLO models coupled with tracking algorithms, assuming linear motion trajectories akin to that observed in CCTV footage. However, such assumptions fall short in virtual meetings, where participant video feed window can abruptly change location across the grid. In an organic video meeting setting, participants frequently join and leave, leading to sudden, non-linear movements on the video grid. This disrupts optical flow-based tracking methods that depend on linear motion. Consequently, standard object detection and tracking methods might mistakenly assign multiple participants to the same tracker. In this paper, we introduce a novel approach to track and re-identify participants in remote video meetings, by utilizing the spatio-temporal priors arising from the data in our domain. This, in turn, increases tracking capabilities compared to the use of general object tracking. Our approach reduces the error rate by 95% on average compared to YOLO-based tracking methods as a baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。