arXiv:2409.11262cs.SDcs.AI2024-09被引 3

构建隐私保护的居家声音数据集,助力老年人健康监测。

The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection

  • 从8位55-80岁老人家中采集7天音频,构建真实居家声景。
  • 用预训练神经网络自动移除语音,保留其他声音事件。
  • 适用于智能养老场景的声音事件检测模型研发。

本文提出一个面向智能家居应用、促进老年人福祉的声音事件检测数据集。在8名年龄55-80岁的参与者家中部署音频采集系统,持续记录7天。通过详细平面图与建筑材料信息记录声学特征,支持人工智能模型部署环境的复现。开发了一种新型自动化语音去除流程,利用预训练音频神经网络检测并剔除含人声片段,同时保留其他声音事件。最终数据集为符合隐私要求的音频记录,准确反映住宅空间内的声音景观与日常活动。论文详述数据集构建方法、采用级联模型架构的语音去除流程,以及语音标签分布分析以验证去语音效果。该数据集可支持专用于家庭环境的声音事件检测模型开发与基准测试。

原文摘要 · Abstract (English)

This paper presents a residential audio dataset to support sound event detection research for smart home applications aimed at promoting wellbeing for older adults. The dataset is constructed by deploying audio recording systems in the homes of 8 participants aged 55-80 years for a 7-day period. Acoustic characteristics are documented through detailed floor plans and construction material information to enable replication of the recording environments for AI model deployment. A novel automated speech removal pipeline is developed, using pre-trained audio neural networks to detect and remove segments containing spoken voice, while preserving segments containing other sound events. The resulting dataset consists of privacy-compliant audio recordings that accurately capture the soundscapes and activities of daily living within residential spaces. The paper details the dataset creation methodology, the speech removal pipeline utilizing cascaded model architectures, and an analysis of the vocal label distribution to validate the speech removal process. This dataset enables the development and benchmarking of sound event detection models tailored specifically for in-home applications.

声音检测智能家居老年健康

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。