arXiv:2508.07863cs.CVcs.LG2025-08被引 7

首个实时可控的视觉语言运动模型,支持精细身体部位控制。

Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model

  • 基于新数据集设计分部位残差量化,实现逐部位精准控制
  • 在500万动作序列上训练,支持长时序生成与未见场景
  • 兼顾实时性与高精度,适合虚拟人、游戏等实际应用

人类动作生成技术具有广泛的应用前景,但现有视觉-语言-运动模型(VLMM)存在可控性瓶颈,体现在对多样化指令响应弱、初始姿态能力有限、长序列表现差、未见场景处理不足以及个体肢体控制粗糙。为此,我们提出Being-M0.5,首个实现实时运行且高度可控的VLMM,性能达当前最优。该模型基于全球最大最全的人类动作数据集HuMo100M,包含超500万自采集动作序列、1亿个跨任务指令实例及细粒度部位标注。引入新型分部位残差量化技术,实现生成过程中的精细化肢体控制。大量实验验证其在多类动作基准上的卓越表现,全面效率分析证实其具备实时生成能力。研究还提供设计洞察与计算分析,推动实用化动作生成器发展。我们认为HuMo100M与Being-M0.5是迈向真实应用的关键进展。

原文摘要 · Abstract (English)

Human motion generation has emerged as a critical technology with transformative potential for real-world applications. However, existing vision-language-motion models (VLMMs) face significant limitations that hinder their practical deployment. We identify controllability as a main bottleneck, manifesting in five key aspects: inadequate response to diverse human commands, limited pose initialization capabilities, poor performance on long-term sequences, insufficient handling of unseen scenarios, and lack of fine-grained control over individual body parts. To overcome these limitations, we present Being-M0.5, the first real-time, controllable VLMM that achieves state-of-the-art performance across multiple motion generation tasks. Our approach is built upon HuMo100M, the largest and most comprehensive human motion dataset to date, comprising over 5 million self-collected motion sequences, 100 million multi-task instructional instances, and detailed part-level annotations that address a critical gap in existing datasets. We introduce a novel part-aware residual quantization technique for motion tokenization that enables precise, granular control over individual body parts during generation. Extensive experimental validation demonstrates Being-M0.5's superior performance across diverse motion benchmarks, while comprehensive efficiency analysis confirms its real-time capabilities. Our contributions include design insights and detailed computational analysis to guide future development of practical motion generators. We believe that HuMo100M and Being-M0.5 represent significant advances that will accelerate the adoption of motion generation technologies in real-world applications. The project page is available at https://beingbeyond.github.io/Being-M0.5.

动作生成视觉语言实时控制分部位建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。