语音情绪与文字内容结合,能更准识别财报会中的回避行为
A Shaky Voice Is Not Always a Dodge: Benchmarking Textual and Vocal Evasion Detection in Earnings Calls

- 构建双模态评测基准,同时标注文字回避和语音自信度
- 现有模型难捕捉不自信语音,尤其在非自信回答上表现差
- 适合研究语音分析、金融对话或人机交互的学者参考
现有财报会议回避检测方法仅依赖文字转录,将回避视为单一维度现象。我们指出,口头交流中的回避本质上是多维的:除言语内容外,表达方式本身携带独立且互补的信息。为此,我们提出DualEvasion,一个面向财报会议问答环节的文字与音频联合避答检测基准。该基准包含60场财报会议中的505个问答对,每对均有两项独立标注:文字回避(直接 vs. 回避)与语音信心度(自信 vs. 不自信)。实验表明,当前最先进的多模态模型在检测语音信心方面表现不佳,尤其在不自信回应上。分析显示这些模型将声学线索孤立解读,而非相对于发言者自身基线进行判断。提供说话人级别的参考仅带来小幅提升,与人类表现仍有显著差距。
原文摘要 · Abstract (English)
Existing approaches to evasion detection in earnings calls focus on textual transcripts, treating evasion as a single-dimensional phenomenon. We argue that evasion in spoken communication is inherently multidimensional: beyond what executives say, how they say it carries independent and complementary information. To study these dimensions jointly, we introduce DualEvasion, a benchmark for evasion detection across text and audio in earnings call Q&A. The benchmark contains 505 annotated question-answer pairs from 60 earnings calls, each with two independent labels: textual evasion (direct vs. evasive) and vocal cues operationalized as speaker confidence (confident vs. unconfident). Our experiments show that state-of-the-art multimodal models struggle to detect vocal confidence, particularly on unconfident responses. Our analysis suggests these models interpret acoustic cues in isolation rather than relative to each speaker's baseline. Providing speaker-level references yields modest improvements, but a substantial gap with human performance remains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。