arXiv:2606.03868cs.CV2026-06

统一建模视频与手部动作去噪,实现灵巧操作生成与数据合成。

Unified Video-Action Joint Denoising for Dexterous Action and Data Generation

论文配图:Unified Video-Action Joint Denoising for Dexterous Action and Data Generation
图 1 · 摘自论文原文
  • 构建视频与手部轨迹联合去噪模型,支持多条件生成。
  • 在动作、视频和纯文本生成任务中均提升精度与流畅性。
  • 适合机器人灵巧操作模拟与高质量数据生成场景。

近期世界动作模型通过将视觉动态先验与可执行机器人动作对齐来构建。本文从分布视角重新审视这一对齐机制。现有方法通常将先验压缩为基于观测的动作策略分布,而本文通过在多种条件设置下建模交互视频与可执行手部轨迹的联合空间,保持分布更广。提出Donk模型,一种用于灵巧手部的统一视频-动作去噪模型。在语言、初始图像和初始手部状态条件下,可采样未来视频与双臂MANO轨迹作为动作策略;无图像条件时,同一去噪架构可从文本条件分布中采样成对的视频-动作轨迹,将对齐的视频先验转化为数据生成引擎。在动作、视频和纯文本生成评估中,Donk均提升了灵巧轨迹准确性,保持强视频保真度,并生成平滑的文本条件动作轨迹,且采用统一训练方案。

原文摘要 · Abstract (English)

Recent world action models leverage video foundation models by aligning broad visual-dynamics priors with executable robot actions. We revisit this alignment from a distributional perspective. Existing formulations typically narrow the aligned prior into an observation-conditioned policy distribution over future actions. In contrast, we keep the distribution broader by modeling the joint space of interaction videos and executable hand trajectories under multiple conditioning regimes. We propose Donk, a unified video-action denoising model for dexterous hands. With language, an initial image, and the initial hand state, Donk samples future videos and bimanual MANO trajectories as an action policy. Without the image condition, the same denoising architecture samples paired video-action rollouts from a text-conditioned distribution, turning the aligned video prior into a data engine. Across action, video, and text-only generation evaluations, Donk improves dexterous trajectory accuracy, preserves strong video fidelity, and produces smooth text-conditioned action rollouts under the same unified training recipe.

灵巧操作视频生成去噪模型数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。