评估开源语音工具在自闭症青少年中的可靠性,发现特征差异影响模型公平性。
Can We Trust Machine Learning? The Reliability of Features from Open-Source Speech Analysis Tools for Speech Modeling
- 对比OpenSMILE与Praat提取的语音特征,分析其一致性
- 不同工具间特征差异显著,影响跨群体模型表现
- 呼吁临床应用中进行领域相关验证以减少偏差
基于机器学习的行为模型依赖于从音视频记录中提取的特征。这些记录通常通过开源工具处理,生成用于分类模型的语音特征。然而,这些工具往往缺乏验证,无法确保其能可靠捕捉行为相关的信息。这一缺口引发了在不同人群和情境下模型可复现性与公平性的担忧。当语音处理工具被用于设计之外的场景时,可能无法公平地捕获行为变异,进而导致偏差。本文评估了两种广泛应用的语音分析工具——OpenSMILE与Praat——在自闭症青少年群体中提取的语音特征的可靠性。结果显示,不同工具间的特征存在显著差异,这种差异影响了模型在不同语境和人口群体中的性能表现。研究建议在临床应用中引入领域相关的验证机制,以提升机器学习模型的可靠性。
原文摘要 · Abstract (English)
Machine learning-based behavioral models rely on features extracted from audio-visual recordings. The recordings are processed using open-source tools to extract speech features for classification models. These tools often lack validation to ensure reliability in capturing behaviorally relevant information. This gap raises concerns about reproducibility and fairness across diverse populations and contexts. Speech processing tools, when used outside of their design context, can fail to capture behavioral variations equitably and can then contribute to bias. We evaluate speech features extracted from two widely used speech analysis tools, OpenSMILE and Praat, to assess their reliability when considering adolescents with autism. We observed considerable variation in features across tools, which influenced model performance across context and demographic groups. We encourage domain-relevant verification to enhance the reliability of machine learning models in clinical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。