arXiv:2510.02287cs.CV2025-10被引 3

用多感官信号提升视频生成的精细控制能力

MultiModal Action Conditioned Video Generation

  • 融合本体感觉、触觉等多模态信号实现精准动作控制
  • 相比纯文本条件模型,生成更准确且时序更稳定
  • 适合需要精细操作的机器人仿真与真实场景应用

当前视频生成模型因缺乏精细控制而难以作为世界模型。通用家用机器人需实时执行精细动作以应对复杂任务和紧急情况。本文引入细粒度多模态动作,涵盖本体感觉、运动觉、力触觉和肌电激活等感知模态,自然支持文本生成模型难以模拟的精细交互。为有效建模多感官细粒度动作,我们提出一种特征学习范式,对齐多模态信息并保留各模态独特性;进一步设计正则化方案,增强动作轨迹特征对复杂交互动态的因果表征能力。实验表明,融合多模态感知显著提升仿真精度并减少时间漂移。大量消融实验与下游应用验证了方法的有效性与实用性。

原文摘要 · Abstract (English)

Current video models fail as world model as they lack fine-graiend control. General-purpose household robots require real-time fine motor control to handle delicate tasks and urgent situations. In this work, we introduce fine-grained multimodal actions to capture such precise control. We consider senses of proprioception, kinesthesia, force haptics, and muscle activation. Such multimodal senses naturally enables fine-grained interactions that are difficult to simulate with text-conditioned generative models. To effectively simulate fine-grained multisensory actions, we develop a feature learning paradigm that aligns these modalities while preserving the unique information each modality provides. We further propose a regularization scheme to enhance causality of the action trajectory features in representing intricate interaction dynamics. Experiments show that incorporating multimodal senses improves simulation accuracy and reduces temporal drift. Extensive ablation studies and downstream applications demonstrate the effectiveness and practicality of our work.

视频生成多模态机器人控制精细动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。