用耳戴设备自动识别日常自言自语,助力心理状态监测
Enabling Automatic Self-Talk Detection via Earables
- 分层分类架构融合声学、语言和上下文信息,逐步分析自言自语
- 在31.1小时数据上达到0.84的宏平均F1值,优于传统方法
- 适合心理健康研究、可穿戴设备开发者及人机交互领域
自言自语——一种可无声或出声进行的内心对话——在情绪调节、认知处理和动机激发中起关键作用,但在日常生活中仍难以被观察和测量。本文提出MutterMeter,一个基于耳戴麦克风音频的移动系统,可在真实场景中自动检测出声自言自语。由于自言自语具有多样的声学特征、语义与语法不完整、出现模式不规则等特性,与传统语音理解模型假设显著不同,检测难度大。MutterMeter采用分层分类架构,通过顺序处理流程逐步整合声学、语言和上下文信息,自适应平衡准确率与计算效率。我们构建并评估了首个此类数据集,包含25名参与者共31.1小时的音频。实验结果表明,MutterMeter在宏平均F1上达到0.84,优于基于LLM和语音情感识别的常规方法。
原文摘要 · Abstract (English)
Self-talk-an internal dialogue that can occur silently or be spoken aloud-plays a crucial role in emotional regulation, cognitive processing, and motivation, yet has remained largely invisible and unmeasurable in everyday life. In this paper, we present MutterMeter, a mobile system that automatically detects vocalized self-talk from audio captured by earable microphones in real-world settings. Detecting self-talk is technically challenging due to its diverse acoustic forms, semantic and grammatical incompleteness, and irregular occurrence patterns, which differ fundamentally from assumptions underlying conventional speech understanding models. To address these challenges, MutterMeter employs a hierarchical classification architecture that progressively integrates acoustic, linguistic, and contextual information through a sequential processing pipeline, adaptively balancing accuracy and computational efficiency. We build and evaluate MutterMeter using a first-of-its-kind dataset comprising 31.1 hours of audio collected from 25 participants. Experimental results demonstrate that MutterMeter achieves robust performance with a macro-averaged F1 score of 0.84, outperforming conventional approaches, including LLM-based and speech emotion recognition models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。