arXiv:2608.07376cs.HCcs.SD2026-08

让AI在合唱中主动参与协调,而非被动回应。

Beyond Call and Response: Modelling Reciprocal Coordination in Human-AI Vocal Ensembles

论文配图:Beyond Call and Response: Modelling Reciprocal Coordination in Human-AI Vocal Ensembles
图 1 · 摘自论文原文
  • 将人机合唱视为耦合动态系统,设计可主动参与的声乐代理架构。
  • 在非均等节拍曲目中实现多对多协同,突破传统节奏网格限制。
  • 适合研究音乐协作、人机交互与集体智能的学者与开发者。

人类与AI的音乐互动通常表现为响应循环:人类演奏后,系统解析并作出回应、伴奏或安排音乐事件。然而,在无指挥的声乐合奏中,演唱者同时持续相互影响,时间和音高均无固定基准,集体组织依赖于多方间的相互调整。本文将此类合奏视为耦合动态系统,提出一种使声乐代理能主动介入而非仅追踪集体状态的研究架构。部分目标曲目具有规整节拍,而另一些则呈现非等时的时间轮廓,无法简化为节拍网格;后者被视作通用框架的难点案例。该架构连接现场多通道采集、方言与演唱风格感知表示、集体状态推断、声乐生成及现场评估。由此形成的新研究议程不仅关注人工歌手能否同步,更探究其存在如何重构人类的协调方式、领导力、风格与音乐传承。

原文摘要 · Abstract (English)

Musical interaction with AI is often organised as a response loop: a human performs, the system interprets that action, and the system answers, accompanies, or schedules a musical event. Unconducted vocal ensembles pose a different problem. Singers act simultaneously and continuously affect one another; neither timing nor pitch is fixed by a conductor, metronome, accompaniment, score, or tuning source. Collective organisation emerges from many-to-many reciprocal adjustment. This paper frames such ensembles as coupled dynamic systems and proposes a research architecture for vocal agents that enter, rather than merely track, their collective states. Some target repertoires are metrical, while others exhibit non-isochronous temporal contours that cannot be reduced to a beat grid; we treat the latter as a hard case for a general framework. The architecture connects multichannel capture in the field to dialect- and singing-aware representation, collective-state inference, vocal generation, and in-situ evaluation. The resulting agenda asks not only whether an artificial singer can synchronise, but how its presence reorganises human coordination, leadership, style, and musical transmission.

人机协作声乐生成集体智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。