arXiv:2609.04958cs.CVcs.RO2026-09

首个直接从第一人称视频生成双手世界坐标轨迹的统一模型。

MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision

论文配图:MINT: A Unified Model for World-Space Camera and Hand Motion Estimation from Scalable Egocentric Pipeline Supervision
图 1 · 摘自论文原文
  • 基于共享时空表示,联合预测相机轨迹与手部状态
  • 在多个基准上提升手部轨迹精度,且推理速度超越标注流水线
  • 适合机器人学习、增强现实等需要精准动作理解的场景

从第一人称视频中恢复世界坐标系下的相机与手部运动是活动理解、机器人学习和增强现实的关键能力。现有系统通常将问题分解为相机运动、深度、手部重建和轨迹优化等多个阶段,导致计算开销大,难以联合建模。本文提出MINT(Minting IN-the-Wild Trajectories),首个能直接从第一人称RGB视频生成完整双臂世界空间轨迹的基础模型。MINT通过单一共享的时空视频表示,联合预测相机轨迹、相机帧内手部状态及每帧手部存在性,再通过显式坐标变换生成世界空间手部运动。由于真实世界坐标标注稀缺,我们开发了开源标注工具EGOPIPELINE,将大量公开第一人称视频转换为结构化相机-手部轨迹监督数据。MINT先在大规模伪标签上预训练,再在少量高质量联合标注数据上微调。在多个公开基准上,MINT实现[xxx]的手部轨迹精度提升、[xxx]的相机轨迹估计改进,且端到端生成速度比标注流水线快[xxx]倍,并可在未见数据集上实现零样本泛化。我们发布模型、代码、标注流水线及一个包含1,021小时的第一人称轨迹数据集。

原文摘要 · Abstract (English)

Recovering camera and hand motion in world coordinates from egocentric video is a key capability for activity understanding, robot learning, and augmented reality. Existing systems typically decompose this problem into separate stages for camera motion, depth, hand reconstruction, and trajectory refinement, resulting in substantial computational overhead and preventing the joint modeling of camera and hand motion. We introduce MINT (Minting IN-the-Wild Trajectories), the first foundation model that directly produces complete world-space two-hand trajectories from ego-centric RGB video. From a single shared spatiotemporal video representation, MINT jointly predicts the camera trajectory, camera-frame hand states, and per-frame hand presence, and then produces world-space hand motion via explicit coordinate transformations. Training such a model at scale is challenging, since paired world-space camera and hand annotations are scarce. We therefore develop an open-source labeling EGOPIPELINE that converts large collections of public egocentric videos into structured camera-and-hand trajectory supervision. MINT is first pretrained on these large-scale pseudo-labels and then fine-tuned on a small set of high-quality joint annotations. Across public benchmarks, MINT achieves [xxx] improvement in world-space hand trajectory accuracy, [xxx] improvement in camera trajectory estimation, and [xxx] faster end-to-end trajectory generation than the labeling pipeline, while generalizing zero-shot to unseen egocentric datasets. We release the model, training and inference code, labeling pipeline, and a curated 1,021-hour egocentric trajectory dataset.

动作估计第一人称视频轨迹生成基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。