研究发现,心理风险预测需至少2000样本训练集才稳定。
Toward Corpus Size Requirements for Training and Evaluating Depression Risk Models Using Spoken Language
- 控制变量下测试不同训练/测试集规模对模型影响
- 训练集少于2000样本时结果不稳定,测试集<1000样本噪声大
- 语言与声学模型表现相似,适合语音心理健康研究者参考
精神健康风险预测是语音领域的重要方向,但多数研究依赖小规模语料。本研究基于超过6.5万条标注数据,在全交叉设计下分析不同训练/测试集规模组合的影响。包含基于语言和声学的两类模型,均采用当前主流方法。同时引入年龄不匹配的测试集。结果表明:(1) 测试集低于1000样本时,即使训练集较大也产生噪声结果;(2) 训练集至少需2000样本才能获得稳定性能;(3) 语言与声学模型在规模变化下行为一致;(4) 年龄不匹配测试集呈现与匹配集相同的模式。此外讨论了标签先验、模型强度、预训练、唯一说话人及数据长度等因素。虽无法给出确切规模标准,但强调未来研究需合理设定训练与测试集规模。
原文摘要 · Abstract (English)
Mental health risk prediction is a growing field in the speech community, but many studies are based on small corpora. This study illustrates how variations in test and train set sizes impact performance in a controlled study. Using a corpus of over 65K labeled data points, results from a fully crossed design of different train/test size combinations are provided. Two model types are included: one based on language and the other on speech acoustics. Both use methods current in this domain. An age-mismatched test set was also included. Results show that (1) test sizes below 1K samples gave noisy results, even for larger training set sizes; (2) training set sizes of at least 2K were needed for stable results; (3) NLP and acoustic models behaved similarly with train/test size variations, and (4) the mismatched test set showed the same patterns as the matched test set. Additional factors are discussed, including label priors, model strength and pre-training, unique speakers, and data lengths. While no single study can specify exact size requirements, results demonstrate the need for appropriately sized train and test sets for future studies of mental health risk prediction from speech and language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。