arXiv:2604.05621cs.CV2026-04被引 3

从第一视角视频重建可交互的3D场景,无需人工标注或固定场景。

FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos

论文配图:FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
图 1 · 摘自论文原文
  • 基于第一视角视频自动发现可动部件并估计运动参数。
  • 在真实与模拟数据上提升50%部分分割指标,运动误差降低5-10倍。
  • 适合需要高精度3D数字孪生的机器人交互与仿真应用。

我们提出FunRec,一种直接从第一视角RGB-D交互视频中重建室内功能性3D数字孪生的方法。不同于依赖受控环境、多状态采集或CAD先验的现有方法,FunRec可直接处理真实世界中的人类交互序列,自动识别可动部件,估计其运动学参数,跟踪3D运动,并在规范空间中重建静态与动态几何结构,生成可用于仿真的网格模型。在新的真实与模拟基准测试中,FunRec显著超越以往方法:部分分割的mIoU最高提升50%,关节与姿态误差降低5-10倍,重建精度明显更高。我们进一步展示了其在URDF/USD导出、手控功能映射及机器人-场景交互中的应用。

原文摘要 · Abstract (English)

We present FunRec, a method for reconstructing functional 3D digital twins of indoor scenes directly from egocentric RGB-D interaction videos. Unlike existing methods on articulated reconstruction, which rely on controlled setups, multi-state captures, or CAD priors, FunRec operates directly on in-the-wild human interaction sequences to recover interactable 3D scenes. It automatically discovers articulated parts, estimates their kinematic parameters, tracks their 3D motion, and reconstructs static and moving geometry in canonical space, yielding simulation-compatible meshes. Across new real and simulated benchmarks, FunRec surpasses prior work by a large margin, achieving up to +50 mIoU improvement in part segmentation, 5-10 times lower articulation and pose errors, and significantly higher reconstruction accuracy. We further demonstrate applications on URDF/USD export for simulation, hand-guided affordance mapping and robot-scene interaction.

3D重建数字孪生机器人交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。