arXiv:2512.24645cs.SD2025-12

AudioFab用工具学习构建智能音频工厂,让非专业人士也能高效处理复杂音频任务。

AudioFab: Building A General and Intelligent Audio Factory through Tool Learning

  • 模块化设计解决工具依赖冲突,简化集成与扩展。
  • 通过智能选型和少样本学习提升复杂任务的效率与准确率。
  • 提供自然语言接口,适合非专家用户快速上手。

当前人工智能正深刻改变音频领域,但众多先进算法与工具仍处于碎片化状态,缺乏统一高效的框架以释放其全部潜力。现有音频智能体框架常因环境配置复杂及工具协作低效而受限。为此,我们提出 AudioFab,一个开源智能音频处理框架,旨在建立开放、智能的音频生态系统。相比现有方案,AudioFab 的模块化设计可化解依赖冲突,简化工具集成与扩展;通过智能工具选择与少样本学习优化工具学习过程,在复杂音频任务中显著提升效率与准确率;同时提供面向非专业用户的友好自然语言接口。作为基础性框架,AudioFab 的核心贡献在于为未来音频与多模态 AI 的研究与开发提供稳定、可扩展的平台。代码已公开于 https://github.com/SmileHnu/AudioFab。

原文摘要 · Abstract (English)

Currently, artificial intelligence is profoundly transforming the audio domain; however, numerous advanced algorithms and tools remain fragmented, lacking a unified and efficient framework to unlock their full potential. Existing audio agent frameworks often suffer from complex environment configurations and inefficient tool collaboration. To address these limitations, we introduce AudioFab, an open-source agent framework aimed at establishing an open and intelligent audio-processing ecosystem. Compared to existing solutions, AudioFab's modular design resolves dependency conflicts, simplifying tool integration and extension. It also optimizes tool learning through intelligent selection and few-shot learning, improving efficiency and accuracy in complex audio tasks. Furthermore, AudioFab provides a user-friendly natural language interface tailored for non-expert users. As a foundational framework, AudioFab's core contribution lies in offering a stable and extensible platform for future research and development in audio and multimodal AI. The code is available at https://github.com/SmileHnu/AudioFab.

音频生成智能体工具学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。