arXiv:2602.17097cs.SD2026-02被引 7

让音频模型像聊天一样理解、生成和编辑复杂声音故事。

AudioChat: Unified Audio Storytelling, Editing, and Understanding with Transfusion Forcing

  • 用大模型模拟用户对话,生成训练数据。
  • 新目标函数实现多轮交互与分步推理。
  • 适合做声音内容创作与智能交互的开发者。

尽管近期取得突破,音频基础模型在处理多源复杂声学场景时仍面临挑战。我们称此类场景为音频故事,通常包含多个说话人及前景/背景音效。相比传统音频任务,音频故事引入了语义、时间与物理层面的新复杂性。为此,我们提出 AudioChat,一个可生成、编辑和理解音频故事的框架。AudioChat 采用基于大模型的工具调用代理,模拟用户与系统间的交互,并以此生成训练数据。我们还提出一种新型音频输血强制(Audio Transfusion Forcing)目标,使模型能同时通过结构化思维链分解高层指令,并执行多轮交互式音频理解与生成。为评估生成与编辑性能,我们设计了三个新指标,直接衡量任务完成度,而非依赖分布评分。强烈推荐访问演示页面以了解 AudioChat 的能力:https://wanchichen.github.io/audiochat/。

原文摘要 · Abstract (English)

Despite recent breakthroughs, audio foundation models struggle in processing complex multi-source acoustic scenes. We refer to this challenging domain as audio stories, which can have multiple speakers and background/foreground sound effects. Compared to traditional audio processing tasks, audio stories introduce new layers of semantic, temporal, and physical complexity. To address this challenge, we propose AudioChat, a framework for developing audio foundation models that can generate, edit, and understand audio stories. AudioChat introduces a new paradigm in which LLM-based toolcalling agents simulate interactions between users and the system, and these simulated dialogues are used as training data. We also introduce a novel Audio Transfusion Forcing objective to train the AudioChat model, allowing it to simultaneously decompose high-level instructions via structured chain-of-thought reasoning and perform interactive multi-turn audio understanding/generation. To evaluate generation and editing performance, we develop three new metrics that directly measure task performance instead of relying upon distribution-based scoring. We highly encourage readers to visit our demo to better understand the capabilities of AudioChat: https://wanchichen.github.io/audiochat/.

音频生成多模态大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。