arXiv:2509.17277eess.AScs.SD2025-09被引 2

构建500个合成提示音数据集,用于人机交互与音频机器学习研究。

BeepBank-500: A Synthetic Earcon Mini-Corpus for UI Sound Research and Psychoacoustics Research

  • 通过参数化生成波形类型、频率、时长等,实现完全可控的合成音设计。
  • 包含300-500段48kHz单声道音频,附带信号特征元数据和轻量基线模型。
  • 适合音频感知、音色分析、触发检测等研究,开源可用且无版权限制。

我们提出BeepBank-500,一个小型、全合成的耳声(earcon)/提示音数据集(300-500个片段),专为快速、无版权顾虑的人机交互与音频机器学习实验而设计。每个音效由参数化配方生成,控制波形类型(正弦、方波、三角波、调频)、基频、时长、振幅包络、振幅调制(AM)及轻量级Schroeder风格混响。使用三种混响设置:干声、'rir small'(小房间)和'rir medium'(中房间)。发布内容包括单声道48kHz WAV音频(16位)、丰富的元数据表(信号/频谱特征),以及两个轻量级基线任务:(i)波形类型分类,(ii)单音基频回归。该数据集适用于耳声分类、音色分析、起始点检测等任务,明确标注许可与局限性。音频采用CC0-1.0公共领域授权;代码为MIT许可。数据集DOI:https://doi.org/10.5281/zenodo.17172015。代码地址:https://github.com/mandip42/earcons-mini-500。

原文摘要 · Abstract (English)

We introduce BeepBank-500, a compact, fully synthetic earcon/alert dataset (300-500 clips) designed for rapid, rights-clean experimentation in human-computer interaction and audio machine learning. Each clip is generated from a parametric recipe controlling waveform family (sine, square, triangle, FM), fundamental frequency, duration, amplitude envelope, amplitude modulation (AM), and lightweight Schroeder-style reverberation. We use three reverberation settings: dry, and two synthetic rooms denoted 'rir small' ('small') and 'rir medium' ('medium') throughout the paper and in the metadata. We release mono 48 kHz WAV audio (16-bit), a rich metadata table (signal/spectral features), and tiny reproducible baselines for (i) waveform-family classification and (ii) f0 regression on single tones. The corpus targets tasks such as earcon classification, timbre analyses, and onset detection, with clearly stated licensing and limitations. Audio is dedicated to the public domain via CC0-1.0; code is under MIT. Data DOI: https://doi.org/10.5281/zenodo.17172015. Code: https://github.com/mandip42/earcons-mini-500.

音频生成人机交互合成数据听觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。