arXiv:2609.04975cs.SD2026-09

首个一步完成的3D空间音频指令编辑框架,支持复杂指令直接生成。

SwanWeave:One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing

论文配图:SwanWeave:One-Stage Multi-Task Instruction-Guided 3D Spatial Audio Editing
图 1 · 摘自论文原文
  • 采用双层路由的专家混合模型,动态组合专家处理复合指令。
  • 在十余种单操作与复合任务上超越现有基线,编辑质量显著提升。
  • 适合需要精准空间音频编辑的研究者与音频内容创作者。

空间音频编辑根据用户指令调整已有声场,同时保留场景其余部分。与传统音频编辑不同,它需联合推理音频事件、空间信息、动态变化和环境信息,处理一阶全向量(FOA)波形。现有语言引导编辑器主要针对常规音频或依赖顺序操作,无法直接支持复杂3D空间指令的一步编辑。我们提出SwanWeave,首个一步式多任务指令引导3D FOA空间音频编辑框架。通过可控房间模拟,从开源语音与音效语料中构建配对的FOA监督数据,覆盖超过十种单操作与复合任务,涵盖四个编辑轴。为应对异构编辑空间,SwanWeave采用空间编辑专家混合(SE-MoE),使用双层路由机制,为复合指令选择任务感知的专家组合,并对帧级编辑决策进行局部路由或空专家处理。我们进一步引入空间偏好优化(SPO),基于直接偏好优化(DPO)设计,包含特定编辑负样本的目标函数,并采用分阶段训练以增强自然语言对齐。实验表明,SwanWeave在所有任务上均优于现有通用音频编辑器与空间音频基线,编辑质量更优。空间音频编辑演示见https://swanaigc.github.io/#swanweave,代码见https://github.com/MM-Speech/SwanWeave。

原文摘要 · Abstract (English)

Spatial audio editing modifies an existing soundfield according to a user's instruction while preserving the rest of the scene. Unlike conventional audio editing, it must reason jointly about audio events, spatial information, dynamic changes, and environmental information in first-order Ambisonic (FOA) waveforms. Existing language-guided editors mainly target conventional audio or rely on sequential operations, and therefore do not directly support one-stage editing for complex 3D spatial instructions. We present SwanWeave, the first one-stage multi-task framework for instruction-guided 3D FOA spatial audio editing. We build paired FOA supervision from open-source speech and sound-effect corpora using controllable room simulation, covering more than ten single-operation and compound tasks across the four editing axes. To handle this heterogeneous edit space, SwanWeave uses Spatial Edit Mixture-of-Experts (SE-MoE) with dual-level routing, selecting task-aware expert combinations for compound instructions and frame-level routed/null experts for local edit decisions. We further introduce Spatial Preference Optimization (SPO), a Direct Preference Optimization (DPO)-based alignment objective with edit-specific negative targets, and adopt staged training to improve natural-language grounding. Experiments show that SwanWeave achieves better editing quality than existing general audio editors and spatial audio baselines across all tasks. Spatial audio editing demos can be found at https://swanaigc.github.io/#swanweave, code can be found at: https://github.com/MM-Speech/SwanWeave.

空间音频指令编辑多任务3D音频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。