无需训练,通过动作参数提升机器人视频生成的逼真度。
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
- 用动作幅度动态调节引导强度,增强运动控制力。
- 调整初始噪声分布,使动作更符合真实动态。
- 适用于需要高精度动作还原的机器人仿真与建模。
从显式动作轨迹生成真实机器人视频是构建有效世界模型和机器人基础模型的关键步骤。我们提出两种无需训练、仅在推理阶段使用的技术,充分挖掘扩散模型中显式动作参数的作用。不再将动作向量视为被动条件信号,而是主动将其用于引导分类器自由引导过程和高斯隐变量的初始化。首先,动作缩放的分类器自由引导根据动作幅度动态调节引导强度,提升对运动强度的控制能力;其次,动作缩放的噪声截断调整初始采样噪声的分布,使其更契合目标运动动态。在真实机器人操作数据集上的实验表明,这些方法显著提升了动作一致性和视觉质量,适用于多种机器人环境。
原文摘要 · Abstract (English)
Generating realistic robot videos from explicit action trajectories is a critical step toward building effective world models and robotics foundation models. We introduce two training-free, inference-time techniques that fully exploit explicit action parameters in diffusion-based robot video generation. Instead of treating action vectors as passive conditioning signals, our methods actively incorporate them to guide both the classifier-free guidance process and the initialization of Gaussian latents. First, action-scaled classifier-free guidance dynamically modulates guidance strength in proportion to action magnitude, enhancing controllability over motion intensity. Second, action-scaled noise truncation adjusts the distribution of initially sampled noise to better align with the desired motion dynamics. Experiments on real robot manipulation datasets demonstrate that these techniques significantly improve action coherence and visual quality across diverse robot environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。