用大模型生成真实场景音频,填补异常检测数据空白
Did You Hear That? Introducing AADG: A Framework for Generating Benchmark Data in Audio Anomaly Detection
- 借大语言模型模拟现实场景,自动生成合成音频
- 支持视频与电话音源等非工业场景,覆盖更广真实环境
- 模块化设计可复用,适合研究音频异常检测的开发者
我们提出一种新型通用音频生成框架,专为异常检测与定位设计。不同于现有主要聚焦工业和机器声音的数据集,该框架拓展至更广泛环境,特别适用于仅有音频可用的真实场景,如视频衍生或电话语音。方法受LLM-Modulo启发,利用大语言模型(LLMs)作为世界模型,模拟现实场景。首先由LLM预测合理情境,再提取其中构成声音、顺序及混合方式,生成连贯音频。每个生成阶段均经严格验证,确保数据可靠性。该框架生成的数据可作为异常检测基准,有助于提升模型在分布外情况下的表现。本工作填补了音频异常检测资源的空白,提供了一种可扩展的多样化真实音频生成工具。
原文摘要 · Abstract (English)
We introduce a novel, general-purpose audio generation framework specifically designed for anomaly detection and localization. Unlike existing datasets that predominantly focus on industrial and machine-related sounds, our framework focuses a broader range of environments, particularly useful in real-world scenarios where only audio data are available, such as in video-derived or telephonic audio. To generate such data, we propose a new method inspired by the LLM-Modulo framework, which leverages large language models(LLMs) as world models to simulate such real-world scenarios. This tool is modular allowing a plug-and-play approach. It operates by first using LLMs to predict plausible real-world scenarios. An LLM further extracts the constituent sounds, the order and the way in which these should be merged to create coherent wholes. Much like the LLM-Modulo framework, we include rigorous verification of each output stage, ensuring the reliability of the generated data. The data produced using the framework serves as a benchmark for anomaly detection applications, potentially enhancing the performance of models trained on audio data, particularly in handling out-of-distribution cases. Our contributions thus fill a critical void in audio anomaly detection resources and provide a scalable tool for generating diverse, realistic audio data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。