无需标注数据,跨语言识别音频中的笑声。
MultiLinguahah : A New Unsupervised Multilingual Acoustic Laughter Segmentation Method

- 将笑声检测转为能量序列异常检测,用BYOL-A编码器提取音频特征
- 在多语言数据集上表现优于现有方法,尤其在非英语场景提升显著
- 适合无标注数据、跨语言语音分析的研究者使用
笑声是跨文化普遍存在的社会性非语言表达,在人际交流中具有促进社交联结与信号传递的重要作用。然而,音频中笑声的检测与分割仍具挑战性,现有机器学习方法多依赖昂贵的人工标注,且数据集主要集中在英语语境。为此,我们提出一种无监督多语言笑声分割方法,将任务建模为基于能量的音频片段异常检测。该方法在由BYOL-A编码器学习的音频表示上应用隔离森林(Isolation Forest)。我们在四个数据集上对比了多种先进笑声检测算法,涵盖单口喜剧、情景喜剧及AudioSet中的通用短音频。结果表明,当前主流方法在多语言环境下性能不佳,而本方法在非英语场景中显著优于现有技术。
原文摘要 · Abstract (English)
Laughter is a social non-vocalization that is universal across cultures and languages, and is crucial for human communication, including social bonding and communication signaling. However, detecting laughter in audio is a challenging task, and segmenting is even more difficult. Currently, Machine Learning methods generally rely on costly manual annotation, and their datasets are mostly based on English contexts. Thus, we propose an unsupervised multilingual method that sets up the laughter segmentation task as an anomaly detection of energy-based segmented audio sequences. Our method applies an Isolation Forest on audio representations learned from BYOL-A encoder. We compare our method with several state-of-the-art laughter detection algorithms on four datasets, including stand-up comedy, sitcoms, and general short audio from AudioSet. Our results show that state-of-the-art methods are not optimized for multilingual contexts, while our method outperforms them in non-English settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。