用开源模型+大模型实现语音监控信息的自动分析与问答检索。
Aud-Sur: An Audio Analyzer Assistant for Audio Surveillance Applications
- 结合开源音频模型提取录音内容,再通过大模型实现自然语言问答。
- 支持用户以提问方式快速获取音频中的关键信息。
- 模块化部署,便于扩展和社区协作开发。
本文提出一种面向多种音频监控应用的音频分析助手工具Aud-Sur(本工作为DEFAME FAKES与EUCINF项目的一部分)。该工具包含两个阶段:音频分析与音频检索。第一阶段利用多个开源音频模型对用户上传的音频文件进行信息提取;第二阶段则基于大语言模型(LLM),支持用户以自然语言问答形式查询已处理音频中的内容。Aud-Sur采用Docker部署,基于微服务架构设计,通过集成开源音频模型、大语言模型及微服务架构,构建了一个高度可扩展、可适配的框架,能够灵活接入更多音频任务,并在音频研究社区中广泛共享与持续开发。
原文摘要 · Abstract (English)
In this paper, we present an audio analyzer assistant tool designed for a wide range of audio-based surveillance applications (This work is a part of our DEFAME FAKES and EUCINF projects). The proposed tool, refered to as Aud-Sur, comprises two main phases Audio Analysis and Audio Retrieval, respectively. In the first phase, multiple open-source audio models are leveraged to extract information from input audio recording uploaded by a user. In the second phase, users interact with the Aud-Sur tool via a natural question-and-answer manner, powered by a large language model (LLM), to retrieve the information extracted from the processed audio file. The Aud-Sur tool was deployed using Docker on a microservices-based architecture design. By leveraging open-source audio models for information extraction, LLM for audio information retrieval, and a microservices-based deployment approach, the proposed Aud-Sur tool offers a highly extensible and adaptable framework that can integrate more audio tasks, and be widely shared within the audio community for further development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。