用可控仿真生成带精细标注的主视角视频,提升交互理解与预测能力。
EgoInteract: Synthetic Egocentric Videos Generation for Interaction Understanding and Anticipation

- 构建可控制的主视角视频生成模拟器,精准调控动作与场景
- 生成带密集时空标注的数据集,支持多种交互任务评估
- 在多个真实数据集上验证,显著提升模型性能,适合视觉研究者
收集大规模带密集空间和时间标注的主视角视频数据集成本高、速度慢,且受限于环境偏见、隐私问题及交互模式覆盖不足。尽管合成数据在多个视觉领域展现出潜力,但在需要时序一致人-物交互的主视角感知中仍研究较少。本文提出EgoInteract,一个可控制的主视角视频生成模拟器,用于建模精细的主视角交互及其时序动态。该模拟器可精确控制相机运动、人体与手部动作、物体操作及场景组合,适用于多样化环境。基于此框架,我们生成了一个带有密集时空标注的合成主视角视频数据集,涵盖时序动作分割、下一活跃物体检测、交互预测和手-物交互检测任务。我们在多个涵盖不同环境、物体类别和交互模式的真实世界主视角基准上评估了使用模拟数据训练的模型,结果表明,在各类任务和数据集上均优于强基线,验证了该仿真方法的有效性与可迁移性。
原文摘要 · Abstract (English)
Collecting large-scale egocentric video datasets with dense spatial and temporal annotations is costly, slow, and often constrained by environmental biases, privacy constraints, and limited coverage of interaction patterns. While synthetic data has shown strong potential in several vision domains, its use for egocentric perception remains relatively underexplored, especially for tasks requiring temporally coherent human-object interactions. In this work, we introduce EgoInteract, a controllable simulator for egocentric video generation designed to model fine-grained egocentric interactions and their temporal dynamics. The simulator enables precise control over camera, human body and hand motion, object manipulation, and scene composition across diverse environments. Building on this framework, we generate a synthetic egocentric video dataset with dense spatial and temporal annotations for temporal action segmentation, next-active object detection, interaction anticipation, and hand-object interaction detection. We evaluate models trained with simulated data on multiple real-world egocentric benchmarks spanning diverse environments, object categories, and interaction patterns. Results show consistent improvements over strong baselines across tasks and datasets, demonstrating the effectiveness and transferability of our simulation-based approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。