arXiv:2604.11038cs.CV2026-04

从第一视角视频中构建可交互3D物体,用功能模板实现精准模拟。

EgoFun3D: Modeling Interactive Objects from Egocentric Videos using Function Templates

  • 用功能模板建模物体部件间动态关系,支持跨平台代码生成。
  • 构建271段真实交互视频数据集,含2D/3D分割与功能标注。
  • 适合做具身智能、物理模拟的开发者和研究者参考。

我们提出EgoFun3D,一个面向第一视角视频中可交互3D物体建模的任务框架、数据集与基准。该任务旨在从日常视频中提取可用于仿真的交互式3D物体。现有工作多关注关节运动,而本研究通过功能模板——一种结构化计算表示——捕捉部件间的通用功能映射(如灶具旋钮旋转影响炉火温度)。功能模板支持精确评估并可直接编译为各仿真平台的可执行代码。为此,我们构建了包含271段第一视角视频的数据集,涵盖复杂现实交互,附带3D几何、2D/3D分割、关节运动及功能模板标注。针对该任务,我们设计四阶段流水线:2D部件分割、三维重建、关节估计与功能模板推断。全面基准测试表明,当前主流方法仍难以胜任,凸显未来研究空间。

原文摘要 · Abstract (English)

We present EgoFun3D, a coordinated task formulation, dataset, and benchmark for modeling interactive 3D objects from egocentric videos. Interactive objects are of high interest for embodied AI but scarce, making modeling from readily available real-world videos valuable. Our task focuses on obtaining simulation-ready interactive 3D objects from egocentric video input. While prior work largely focuses on articulations, we capture general cross-part functional mappings (e.g., rotation of stove knob controls stove burner temperature) through function templates, a structured computational representation. Function templates enable precise evaluation and direct compilation into executable code across simulation platforms. To enable comprehensive benchmarking, we introduce a dataset of 271 egocentric videos featuring challenging real-world interactions with paired 3D geometry, segmentation over 2D and 3D, articulation and function template annotations. To tackle the task, we propose a 4-stage pipeline consisting of: 2D part segmentation, reconstruction, articulation estimation, and function template inference. Comprehensive benchmarking shows that the task is challenging for off-the-shelf methods, highlighting avenues for future work.

3D建模具身智能功能模板第一视角视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。