arXiv:2410.16259cs.CVcs.GR2024-10ICLR被引 3

从手机视频学动物人类行为,构建可交互的虚拟仿真模型

Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos

  • 通过长时间单视角视频,非侵入式追踪3D行为轨迹
  • 构建4D时空表示,实现跨时间持久追踪与行为建模
  • 适用于宠物和人类,无需标记或多摄像头

我们提出Agent-to-Sim(ATS),一种从自然长期视频中学习3D代理交互行为模型的框架。与依赖标记追踪和多视角相机的以往方法不同,ATS通过在单一环境中长期(如一个月)记录的单目RGBD视频,非侵入式地学习动物和人类的自然行为。建模3D行为需长时间持续的3D追踪(如确定某点始终对应同一位置)。为此,我们设计了一种粗到精的配准方法,将代理与摄像机随时间映射至统一3D空间,生成完整且持久的4D时空表示。随后,利用从4D重建中提取的感知-运动配对数据训练代理行为生成模型。该框架实现从真实视频到交互行为模拟器的迁移。我们在猫、狗、兔子及人类上验证了方法,仅使用智能手机拍摄的单目RGBD视频。

原文摘要 · Abstract (English)

We present Agent-to-Sim (ATS), a framework for learning interactive behavior models of 3D agents from casual longitudinal video collections. Different from prior works that rely on marker-based tracking and multiview cameras, ATS learns natural behaviors of animal and human agents non-invasively through video observations recorded over a long time-span (e.g., a month) in a single environment. Modeling 3D behavior of an agent requires persistent 3D tracking (e.g., knowing which point corresponds to which) over a long time period. To obtain such data, we develop a coarse-to-fine registration method that tracks the agent and the camera over time through a canonical 3D space, resulting in a complete and persistent spacetime 4D representation. We then train a generative model of agent behaviors using paired data of perception and motion of an agent queried from the 4D reconstruction. ATS enables real-to-sim transfer from video recordings of an agent to an interactive behavior simulator. We demonstrate results on pets (e.g., cat, dog, bunny) and human given monocular RGBD videos captured by a smartphone.

行为建模视频生成4D重建仿真迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。