arXiv:2606.30266cs.LGcs.AI2026-06中稿 · the Conference on …

让运动语言模型持续学习新动作,不遗忘旧技能。

Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation

论文配图:Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation
图 1 · 摘自论文原文
  • 用低秩适配变体和专家混合架构,避免任务间干扰。
  • 在五个任务中几乎零遗忘,生成与描述质量保持高位。
  • 路由选择硬专家比软融合更有效,适合长期学习场景。

运动-语言代理需同时具备理解人类动作(运动到文本,M2T)和从自然语言生成动作(文本到运动,T2M)的能力。尽管基础模型在静态场景表现良好,但动态环境中自主代理必须持续学习新运动概念(如新体育风格或专用手势),同时避免对已有技能的灾难性遗忘。本文研究在序列任务暴露下双向运动-语言学习中的稳定性-可塑性权衡。基于冻结的大语言模型主干,提出用于缓解任务间干扰的低秩适配(LoRA)变体。特别设计基于自编码器路由的专家混合架构,在推理时选择特定任务专家,无需任务标签。为评估方法,构建源自HumanML3D的可复现五任务基准,通过语义聚类运动描述生成。实验表明,两种方向均实现近零遗忘,同时保持高生成与标注质量。进一步发现,路由引导的硬专家选择显著优于软专家融合,说明保持专家隔离对性能至关重要。最后观察到词元级准确率与下游生成质量之间可能存在分歧,提示未来研究需建立更全面的评估体系。

原文摘要 · Abstract (English)

Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it from natural language (text-to-motion, T2M). While foundational models have achieved strong performance in static settings, autonomous agents operating in dynamic environments must continuously incorporate new motion concepts -- such as novel athletic styles or specialized gestures -- without catastrophic forgetting of previously acquired skills. We investigate the stability-plasticity trade-off in bidirectional motion-language learning under sequential task exposure. Building on a frozen large language model backbone, we introduce low-rank adaptation (LoRA) variants designed to mitigate inter-task interference. We specifically propose mixture-of-experts architectures that utilize an autoencoder-based router to select task-specific experts at inference time, so that no task-label is needed. To evaluate these methods, we establish a reproducible five-task benchmark derived from HumanML3D through semantic clustering of motion descriptions. Our experimental results demonstrate near-zero forgetting across both M2T and T2M directions while maintaining high generation and captioning quality. Furthermore, we show that hard expert selection via routing significantly outperforms soft expert blending in quality metrics, indicating that preserving expert isolation is critical for maintaining performance in our continual learning setting. Finally, we observe that a divergence between token-level accuracy and downstream generation quality may occur, highlighting the need for more comprehensive evaluation protocols in future research on lifelong motion-language agents.

运动生成持续学习LoRA多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。