智能代理StrixAE解决真实场景下音频增强的复杂失真耦合问题。
StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios

- 基于多模态大模型构建智能代理,协同多个增强与个性化模型。
- 两阶段训练提升鲁棒性,实现感知质量、结构一致性和格式有效性的联合优化。
- 适合需要自动、可靠音频修复的工业级应用,如语音助手、会议系统。
真实场景中的音频增强面临复杂的失真耦合问题,且需个性化处理,现有方案难以兼顾二者。为提升系统鲁棒性并实现自主运行,我们提出基于多模态大语言模型(MLLM)的StrixAE智能代理。该代理以MLLM为控制器,协调多个音频增强与个性化模型。通过两阶段训练:首先在AcoustBench上进行思维链监督微调以建立基础推理与工具调用能力;其次采用专为音频恢复流程设计的音频感知强化学习(APRL),联合优化格式有效性、结构连贯性和感知质量。不同于通用强化学习,APRL引入结构化奖励机制,强制执行可执行的处理流程与逻辑顺序,避免幻觉工具生成。在真实世界测试数据集上,本方法优于多数开源及商用方案,在多个感知指标上达到当前最佳性能,并展现出强泛化鲁棒性。
原文摘要 · Abstract (English)
Audio enhancement in real-world scenarios involves complex distortion couplings and requires personalized enhancement. Existing solutions struggle to address both simultaneously. To improve robustness and enable autonomous operation in such scenarios, we propose StrixAE, an agent based on a multimodal large language model (MLLM). StrixAE leverages the MLLM as a controller to coordinate multiple audio enhancement and personalization models. To further enhance system robustness, reduce artifacts, and improve generalization across diverse real-world scenarios, StrixAE is trained through a two-stage process: first, CoT supervised fine-tuning on AcoustBench to ground basic reasoning and tool invocation; second, Audio Perception Reinforcement Learning (APRL), a reward design specifically tailored for audio restoration pipelines that jointly optimizes format validity, structural coherence, and perceptual quality. Unlike generic RL fine-tuning, APRL introduces structured rewards that enforce executable pipelines and logical section ordering, enabling the agent to produce reliable, interpretable enhancement plans without hallucinated tools. Based on real-world test datasets, our proposed method outperforms most existing open-source and proprietary solutions, achieving state-of-the-art performance across multiple perceptual metrics and demonstrating strong generalization robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。