让AI像人一样一步步思考,生成更贴合视频的高质量音频。
ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
- 用思维链分步推理,逐步生成和编辑音频
- 在多个评测中达到当前最佳效果,尤其擅长复杂场景
- 适合音视频创作、AI辅助设计等需要精细控制的场景
尽管端到端视频转音频技术已显著进步,但生成能真实反映视觉细节的高保真音频仍具挑战。这需要对视觉动态、声学环境和时间关系等进行复杂推理。我们提出ThinkSound框架,利用思维链(CoT)推理实现分步、可交互的视频音频生成与编辑。该方法分为三个阶段:基础拟音生成构建语义连贯的声音场景,通过用户精准交互进行对象中心的优化,以及基于自然语言指令的定向编辑。每个阶段均由多模态大模型生成上下文一致的思维链推理,指导统一的音频基础模型。此外,我们构建了AudioCoT数据集,包含结构化推理标注,连接视觉内容、文本描述与声音合成。实验表明,ThinkSound在视频转音频任务中于音频指标和思维链指标上均达领先水平,并在Movie Gen Audio的分布外测试中表现优异。
原文摘要 · Abstract (English)
While end-to-end video-to-audio generation has greatly improved, producing high-fidelity audio that authentically captures the nuances of visual content remains challenging. Like professionals in the creative industries, this generation requires sophisticated reasoning about items such as visual dynamics, acoustic environments, and temporal relationships. We present ThinkSound, a novel framework that leverages Chain-of-Thought (CoT) reasoning to enable stepwise, interactive audio generation and editing for videos. Our approach decomposes the process into three complementary stages: foundational foley generation that creates semantically coherent soundscapes, interactive object-centric refinement through precise user interactions, and targeted editing guided by natural language instructions. At each stage, a multimodal large language model generates contextually aligned CoT reasoning that guides a unified audio foundation model. Furthermore, we introduce AudioCoT, a comprehensive dataset with structured reasoning annotations that establishes connections between visual content, textual descriptions, and sound synthesis. Experiments demonstrate that ThinkSound achieves state-of-the-art performance in video-to-audio generation across both audio metrics and CoT metrics, and excels in the out-of-distribution Movie Gen Audio benchmark. The project page is available at https://ThinkSound-Project.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。