用概率化流模型生成音色一致的虚拟乐器
FlowSynth: Instrument Generation Through Distributional Flow Matching and Test-Time Search
- 用分布流匹配建模预测不确定性,让生成更灵活
- 测试时搜索多轨迹,提升不同音高音量下的音色一致性
- 适合音乐生成与实时演奏场景
虚拟乐器生成需在不同音高和力度下保持音色一致,现有逐音符模型难以解决此问题。我们提出FlowSynth,结合分布流匹配(DFM)与测试时优化,实现高质量乐器合成。不同于传统确定性流匹配,DFM将速度场参数化为高斯分布,通过负对数似然进行优化,使模型能表达预测不确定性。这一概率化设定支持合理测试时搜索:我们采样多个轨迹并按模型置信度加权,选择最大化音色一致性的输出。FlowSynth在单音质量与跨音符一致性上均优于当前最优基线TokenSynth。结果表明,在流匹配中建模预测不确定性,并结合音乐特异性的一致性目标,是实现可实时演奏的专业级虚拟乐器的有效路径。
原文摘要 · Abstract (English)
Virtual instrument generation requires maintaining consistent timbre across different pitches and velocities, a challenge that existing note-level models struggle to address. We present FlowSynth, which combines distributional flow matching (DFM) with test-time optimization for high-quality instrument synthesis. Unlike standard flow matching that learns deterministic mappings, DFM parameterizes the velocity field as a Gaussian distribution and optimizes via negative log-likelihood, enabling the model to express uncertainty in its predictions. This probabilistic formulation allows principled test-time search: we sample multiple trajectories weighted by model confidence and select outputs that maximize timbre consistency. FlowSynth outperforms the current state-of-the-art TokenSynth baseline in both single-note quality and cross-note consistency. Our approach demonstrates that modeling predictive uncertainty in flow matching, combined with music-specific consistency objectives, provides an effective path to professional-quality virtual instruments suitable for real-time performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。