用朗读脚本收集自然语音,兼顾内容控制与隐私安全
Collecting Prosody in the Wild: A Content-Controlled, Privacy-First Smartphone Protocol and Empirical Evaluation
- 用固定脚本朗读,统一内容避免语义干扰
- 在手机端提取特征并立即删除音频,保障隐私
- 适合心理学、语音分析等需真实语音数据的研究
为进行语音韵律分析,日常语音数据采集面临语义与韵律混淆、隐私限制和参与者配合度低的挑战。本文提出并实证评估了一种内容可控、隐私优先的智能手机采集协议:通过脚本化朗读句子标准化词汇内容(包括提示情绪倾向),同时捕捉自然的韵律表达差异。该协议在设备端完成韵律特征提取,立即删除原始音频,仅传输衍生特征用于分析。我们在大规模研究中部署该协议(N=560;9,877次录音),评估了参与率与数据质量,并基于提取特征进行了诊断性预测任务,成功预测了说话人性别及瞬时情绪状态(正负性、唤醒度)。讨论了该协议对后续研究的应用前景与改进方向。
原文摘要 · Abstract (English)
Collecting everyday speech data for prosodic analysis is challenging due to the confounding of prosody and semantics, privacy constraints, and participant compliance. We introduce and empirically evaluate a content-controlled, privacy-first smartphone protocol that uses scripted read-aloud sentences to standardize lexical content (including prompt valence) while capturing naturalistic variation in prosodic delivery. The protocol performs on-device prosodic feature extraction, deletes raw audio immediately, and transmits only derived features for analysis. We deployed the protocol in a large study (N = 560; 9,877 recordings), evaluated compliance and data quality, and conducted diagnostic prediction tasks on the extracted features, predicting self-reported speaker sex and momentary affective states (valence, arousal). We discuss implications and directions for advancing and deploying the protocol.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。