arXiv:2608.04176cs.SD2026-08被引 1

用手机麦克风实时检测恐慌声音,误报率低且无需用户配合。

Smartphone Audio Based Distress Detection

论文配图:Smartphone Audio Based Distress Detection
图 1 · 摘自论文原文
  • 分两阶段的SVM学习框架,识别哭喊等恐惧语音。
  • 在复杂环境仍保持高检出率,误报率相当于每3-4小时1次社交帖子。
  • 适合日常安全监控,尤其适用于无人值守的紧急预警场景。

我们研究了一种无感、全天候的恐慌检测与报警系统Always Alert,其只需手机保持活跃,无需用户主动参与。该系统利用智能手机内置麦克风(所有手机均具备)和数据网络,提出一种新型两阶段监督学习框架,基于支持向量机(SVM),在手机端实时监测自然的恐惧发声——本研究中为尖叫与哭泣。核心挑战是在保证高恐慌检出率的同时,将误报率控制在可接受范围,使普通用户日常生活不受干扰。通过在多种环境背景下精心筛选的恐慌音频指纹训练模型,优化检出率与误报率(FAR)。实验验证了该框架在复杂音频环境中仍具优异性能;进一步利用误报的时间连续性特征,有效降低误报率。通过志愿者在日常生活中持续采集的多小时音频数据测试,证明该框架可在任何时间、地点可靠运行。最终实现的平均误报开销相当于每3至4小时一次社交媒体发帖。

原文摘要 · Abstract (English)

We investigate an unobtrusive and $24\times7$ human distress detection and signaling system, Always Alert, that requires the smartphone, and not its human owner, to be on alert. The system leverages the microphone sensor, at least one of which is available on every phone, and assumes the availability of a data network. We propose a novel two-stage supervised learning framework, using support vector machines (SVMs), that executes on a user's smartphone and monitors natural vocal expressions of fear---screaming and crying in our study---when a human being is in harm's way. The challenge is to achieve a high distress detection rate while ensuring that the false alarm rate is a manageable overhead, while a typical smartphone user goes about living life as usual. We train the learning framework with carefully selected audio fingerprints of distress and of varied environmental contexts. The audio is used to tune the learning framework to obtain a desirable distress detection rate and false alarm rate (FAR). The ability of the proposed framework to detect distress in rather challenging audio environments is demonstrated. Exploiting the time contiguous nature of false alarms further allows us to reduce the FAR. We show the feasibility of using our framework anytime and anywhere by testing it over many hours of audio fingerprints recorded by volunteers on their smartphones, as they went about their daily routines. We are able to achieve high distress detection rates at an average overhead that is equivalent to about 1 facebook post every 3 to 4 hours.

音频检测安全预警隐私保护智能手机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。