用自监督方法构建可现场演奏的神经合成器,让音乐家直接操控声音生成过程。
Architecture and Affordances of PLAUD: Performative Latents and Unsupervised DDSP

- 基于变分DDSP与潜空间平滑,结合多尺度谱损失和对抗训练
- 支持实时音色变形与轨迹采样,控制响应延迟低于20毫秒
- 适合电子音乐表演者探索非传统音色,关注设计如何塑造演奏行为
PLAUD(Performative Latents and Unsupervised DDSP)是一款基于NoiseBandNet、在小规模个人声学语料上训练的神经合成器及Max for Live乐器,用于现场电子音乐创作。系统融合变分DDSP合成模型、潜空间平滑、多尺度谱损失与对抗损失,并可选使用Transformer先验;通过组件限制、波形整形和先验反馈等直接干预合成链的操作实现音色动态调整。Max for Live界面提供控制生成、轨迹采样与调制三种交互模式。文中贯穿一种具身性分析,指出系统的表演特性源于架构设计本身,而非后期附加。论文既提供了系统的技术实现细节,也从情境化视角剖析其在实时电子音乐表演中的可用性与意义。
原文摘要 · Abstract (English)
PLAUD (Performative Latents and Unsupervised DDSP) is a neural synthesizer and Max for Live instrument for live electronic music, built on NoiseBandNet and trained on small personal sound corpora. We present its architecture, combining a variational DDSP synthesis model, latent smoothing, multi-scale spectral and adversarial losses, and an optional transformer prior, alongside a set of bending operations that intervene directly in the synthesis chain: component limiting, waveshaping, and prior feedback. The Max for Live interface exposes control generation, trajectory sampling, and modulation as primary modes of interaction. Throughout, we thread an affordance analysis arguing that the system's performative character follows from architectural decisions rather than being designed on top of them. The paper contributes both a technical account of the system and a situated affordance analysis of its role in live electronic music performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。