高训练度演讲者在真实会议中语音特征反致金融风险预测失效
The Acoustic Camouflage Phenomenon: Re-evaluating Speech Features for Financial Risk Prediction

- 用双流融合架构对比语音与文本模型表现
- 加入语音特征后风险事件召回率从66.25%降至47.08%
- 发现媒体训练导致的语音伪装效应破坏多模态学习
在计算副语言学中,从语音信号检测认知负荷与欺骗是重要研究方向。近期工作尝试将此类声学框架应用于企业财报电话会议,以预测股市剧烈波动。本研究实证考察了在真实远程会议环境中,对高度训练过的发言者使用声学特征提取(音高、抖动、停顿)的局限性。采用双流晚期融合架构,对比基于声学的流与基准自然语言处理(NLP)流。孤立的NLP模型对尾部风险下行事件的召回率达66.25%。令人意外的是,通过晚期融合加入声学特征后性能显著下降,召回率降至47.08%。我们识别出该现象为‘语音伪装’(Acoustic Camouflage),即媒体训练导致的语音调控引入矛盾噪声,干扰多模态元学习器。本研究提出语音处理在高风险金融预测中的边界条件。
原文摘要 · Abstract (English)
In computational paralinguistics, detecting cognitive load and deception from speech signals is a heavily researched domain. Recent efforts have attempted to apply these acoustic frameworks to corporate earnings calls to predict catastrophic stock market volatility. In this study, we empirically investigate the limits of acoustic feature extraction (pitch, jitter, and hesitation) when applied to highly trained speakers in in-the-wild teleconference environments. Utilizing a two-stream late-fusion architecture, we contrast an acoustic-based stream with a baseline Natural Language Processing (NLP) stream. The isolated NLP model achieved a recall of 66.25% for tail-risk downside events. Surprisingly, integrating acoustic features via late fusion significantly degraded performance, reducing recall to 47.08%. We identify this degradation as Acoustic Camouflage, where media-trained vocal regulation introduces contradictory noise that disrupts multimodal meta-learners. We present these findings as a boundary condition for speech processing applications in high-stakes financial forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。