arXiv:2509.03529cs.CLcs.AI2025-09

用多模态信息构建财报电话会的分层语篇结构,提升跨评估能力。

Multimodal Proposal for an AI-Based Tool to Increase Cross-Assessment of Messages

  • 将财报通话建模为包含问答对的分层语篇树,融合文本、语音、视频与元数据
  • 两阶段Transformer生成稳定且语义丰富的全局嵌入,反映情感、逻辑与主题一致性
  • 适用于金融之外的医疗、教育、政治等高风险非结构化对话场景

财报电话会是兼具脚本化管理层发言与非脚本化分析师互动的丰富金融沟通源。尽管近年金融情绪分析已整合文本与声调等多模态信号,但多数系统仍依赖文档级或句子级模型,难以捕捉此类交互的分层话语结构。本文提出一种新型多模态框架,通过将财报电话会编码为分层话语树来生成语义丰富且结构感知的嵌入表示。每个节点(单人发言或问答对)融合文本、音频、视频中的情感信号及连贯性得分、主题标签、回答覆盖度等结构化元数据。采用两阶段Transformer架构:第一阶段在节点层面利用对比学习编码多模态内容与话语元数据;第二阶段合成整个会议的全局嵌入。实验表明,生成的嵌入具有稳定性与语义意义,能准确反映情感基调、结构逻辑与主题对齐。该方法不仅适用于金融预测与话语评估,还可推广至远程医疗、教育、政治话语等高风险非结构化沟通领域,提供可解释性强的多模态话语表征通用方案。

原文摘要 · Abstract (English)

Earnings calls represent a uniquely rich and semi-structured source of financial communication, blending scripted managerial commentary with unscripted analyst dialogue. Although recent advances in financial sentiment analysis have integrated multi-modal signals, such as textual content and vocal tone, most systems rely on flat document-level or sentence-level models, failing to capture the layered discourse structure of these interactions. This paper introduces a novel multi-modal framework designed to generate semantically rich and structurally aware embeddings of earnings calls, by encoding them as hierarchical discourse trees. Each node, comprising either a monologue or a question-answer pair, is enriched with emotional signals derived from text, audio, and video, as well as structured metadata including coherence scores, topic labels, and answer coverage assessments. A two-stage transformer architecture is proposed: the first encodes multi-modal content and discourse metadata at the node level using contrastive learning, while the second synthesizes a global embedding for the entire conference. Experimental results reveal that the resulting embeddings form stable, semantically meaningful representations that reflect affective tone, structural logic, and thematic alignment. Beyond financial reporting, the proposed system generalizes to other high-stakes unscripted communicative domains such as tele-medicine, education, and political discourse, offering a robust and explainable approach to multi-modal discourse representation. This approach offers practical utility for downstream tasks such as financial forecasting and discourse evaluation, while also providing a generalizable method applicable to other domains involving high-stakes communication.

多模态话语结构财报分析嵌入表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。