arXiv:2412.00591cs.SDcs.AI2024-12被引 1

用可视化工具探索音频数据,快速发现模式与异常。

Audio Atlas: Visualizing and Exploring Audio Datasets

  • 通过文本-音频嵌入将声音映射到二维空间
  • 支持高效语义搜索和动态可视化分析
  • 开源可扩展,适合音频研究与数据探索

我们提出 Audio Atlas,一个基于文本-音频嵌入的交互式网页应用,用于可视化音频数据。该系统采用对比嵌入模型和向量数据库,实现高效的数据管理与语义搜索。通过将音频嵌入映射至二维空间,并利用 DeepScatter 实现动态可视化。Audio Atlas 具备良好的可扩展性,支持便捷集成新数据集,帮助用户深入理解音频数据、识别模式与异常。我们已开源代码库,并提供包含多种音频与音乐数据集的初始版本。

原文摘要 · Abstract (English)

We introduce Audio Atlas, an interactive web application for visualizing audio data using text-audio embeddings. Audio Atlas is designed to facilitate the exploration and analysis of audio datasets using a contrastive embedding model and a vector database for efficient data management and semantic search. The system maps audio embeddings into a two-dimensional space and leverages DeepScatter for dynamic visualization. Designed for extensibility, Audio Atlas allows easy integration of new datasets, enabling users to better understand their audio data and identify both patterns and outliers. We open-source the codebase of Audio Atlas, and provide an initial implementation containing various audio and music datasets.

音频可视化嵌入模型数据探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。