arXiv:2606.19914cs.ROcs.AI2026-06

人机协作生成音乐,机器人能听懂意图并实时演奏。

Co-policy: Responsive Human-Robot Co-Creation for Musical Performances

论文配图:Co-policy: Responsive Human-Robot Co-Creation for Musical Performances
图 1 · 摘自论文原文
  • 用语义锚点+视觉语言模型生成共创计划
  • 单次前向传播实现低延迟多模态动作输出
  • 实机击打乐器实验验证响应更准更频繁

艺术长久以来是人类创造力的重要表达。具身人工智能为生成模型通过物理动作参与创造性活动提供了新路径,而非仅限于数字内容生成。在人机音乐共创中,如何将音乐语义理解与实时、可执行的表演相连接仍具挑战。本文提出 Co-policy 框架,将语义意图定位、受约束的音乐变化和视觉-运动执行三者分离处理。为实现语义定位,Co-policy 使用预推理语义锚点与微调后的 Qwen-vl 计划器(F-Qwen),将语音、现场音乐种子及视觉观察转化为结构化共创计划。为支持低延迟执行,引入高斯混合视觉-运动策略(GMP),作为条件混合密度策略,单次前向传播即可将目标音符与视觉上下文映射为多模态机器人动作。与仅复现用户指定音符的播放系统不同,Co-policy 在音乐与物理约束下生成互补性音乐回应。真实机器人击打乐器实验、消融研究及专家评估表明,相比扩散策略与简化基线,其意图对齐度、执行准确率与响应频率均显著提升,验证了具身动作生成在人机协同创作中的关键作用。

原文摘要 · Abstract (English)

Art has long stood as a pivotal expression of human creativity. Embodied artificial intelligence offers a route for generative models to participate in that creativity through physical action rather than disembodied digital content. In robotic music co-creation, it is challenging to connect semantic musical understanding with real-time and physically executable performance. We present Co-policy, a framework for human-robot musical co-creation that separates semantic intent grounding, constrained musical variation, and visuomotor execution. To ground musical semantics, Co-policy uses pre-inference semantic anchors and a fine-tuned Qwen-vl planner (F-Qwen) to transform speech, live musical seeds, and visual observations into structured co-creation plans. To support low-latency execution, Co-policy introduces a Gaussian-Mixture Visuomotor Policy (GMP), implemented as a conditional mixture-density policy that maps target notes and visual context to multimodal robot actions in a single forward pass. Unlike robotic playback systems that merely reproduce user-specified notes, Co-policy generates complementary musical responses under both musical and physical constraints. Real-robot chime experiments, ablations, and expert evaluation show improved intent alignment, execution accuracy, and response frequency over diffusion-policy and ablated baselines, supporting physically grounded action generation as a key requirement for embodied human-AI co-creation.

人机共创具身智能音乐生成机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。