用声音和文字分析财报电话会,预测市场波动
The Sound of Risk: A Multimodal Physics-Informed Acoustic Model for Forecasting Market Volatility and Enhancing Market Interpretability
- 融合语音声学与文本情绪,构建可解释的情绪空间
- 预测30天市场波动率方差达43.8%,关键在讲话切换时的情绪变化
- 适合投资者和监管者追踪企业隐性风险
金融市场的信息不对称常因企业精心构建的叙事而加剧,削弱了传统文本分析的效果。本文提出一种新型多模态框架,将财报电话会中的文本情感与高管发声时的声带动力学特征结合。核心是物理信息声学模型(PIAM),通过非线性声学方法在信号截断等失真条件下稳健提取情绪信号。声学与文本情绪共同映射到可解释的三维情感状态空间(张力、稳定、唤醒)。基于1,795场财报电话会(约1,800小时)数据,构建了高管从正式陈述转向即兴问答过程中情绪动态变化的特征。关键发现:虽无法预测股价方向,但多模态特征可解释高达43.8%的30天实际波动率方差。波动率预测主要由高管从脚本转为自由发言时的情绪变化驱动,尤其是财务官的文本稳定性下降与声学不稳定性上升,以及首席执行官的唤醒度波动。消融实验表明,该多模态方法显著优于仅使用财务数据的基线模型,验证了声学与文本模态的互补性。通过解码来自可验证生物信号的不确定性指标,该方法为投资者与监管机构提供了增强市场可解释性、识别隐藏企业不确定性的有力工具。
原文摘要 · Abstract (English)
Information asymmetry in financial markets, often amplified by strategically crafted corporate narratives, undermines the effectiveness of conventional textual analysis. We propose a novel multimodal framework for financial risk assessment that integrates textual sentiment with paralinguistic cues derived from executive vocal tract dynamics in earnings calls. Central to this framework is the Physics-Informed Acoustic Model (PIAM), which applies nonlinear acoustics to robustly extract emotional signatures from raw teleconference sound subject to distortions such as signal clipping. Both acoustic and textual emotional states are projected onto an interpretable three-dimensional Affective State Label (ASL) space-Tension, Stability, and Arousal. Using a dataset of 1,795 earnings calls (approximately 1,800 hours), we construct features capturing dynamic shifts in executive affect between scripted presentation and spontaneous Q&A exchanges. Our key finding reveals a pronounced divergence in predictive capacity: while multimodal features do not forecast directional stock returns, they explain up to 43.8% of the out-of-sample variance in 30-day realized volatility. Importantly, volatility predictions are strongly driven by emotional dynamics during executive transitions from scripted to spontaneous speech, particularly reduced textual stability and heightened acoustic instability from CFOs, and significant arousal variability from CEOs. An ablation study confirms that our multimodal approach substantially outperforms a financials-only baseline, underscoring the complementary contributions of acoustic and textual modalities. By decoding latent markers of uncertainty from verifiable biometric signals, our methodology provides investors and regulators a powerful tool for enhancing market interpretability and identifying hidden corporate uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。