测试大模型在财报电话会中的隐含情绪识别能力,发现其与股价走势关联性不足。
Can AI Read Between The Lines? Benchmarking LLMs On Financial Nuance
- 用财报电话录音文本测试主流大模型的情绪判断能力
- 模型预测情绪与股价变动相关性低于人类分析师水平
- 适合金融风控、投资决策辅助等需要精准情绪分析的场景
截至2025年,生成式人工智能已成为各行业提效的核心工具。除文本生成外,其在编程、数据分析和研究流程中也扮演关键角色。随着大语言模型(LLMs)持续演进,评估其输出在专业高风险领域(如金融)中的可靠性与准确性至关重要。当前多数LLM将文本转换为数值向量,用于余弦相似度搜索等操作,但这一抽象过程可能导致对情感基调的误判,尤其在金融语境下的细微表达中更为明显。尽管大模型在日常语言情绪识别上表现良好,但在财报电话会中常见的委婉表述、前瞻性语言及行业术语构成的复杂语境下,往往难以准确解析。本文基于圣克拉拉大学微软实践项目(由查理·戈尔德伯格教授主导),对微软Copilot、OpenAI ChatGPT、谷歌Gemini及传统机器学习模型在金融文本情绪分析上的表现进行基准测试。使用微软财报电话会文本,评估大模型推断情绪与市场情绪及股价波动的相关性,并检验提示工程对结果的改进效果。通过可视化手段分析情绪一致性,考察各业务线情绪趋势对整体股价的影响程度。
原文摘要 · Abstract (English)
As of 2025, Generative Artificial Intelligence (GenAI) has become a central tool for productivity across industries. Beyond text generation, GenAI now plays a critical role in coding, data analysis, and research workflows. As large language models (LLMs) continue to evolve, it is essential to assess the reliability and accuracy of their outputs, especially in specialized, high-stakes domains like finance. Most modern LLMs transform text into numerical vectors, which are used in operations such as cosine similarity searches to generate responses. However, this abstraction process can lead to misinterpretation of emotional tone, particularly in nuanced financial contexts. While LLMs generally excel at identifying sentiment in everyday language, these models often struggle with the nuanced, strategically ambiguous language found in earnings call transcripts. Financial disclosures frequently embed sentiment in hedged statements, forward-looking language, and industry-specific jargon, making it difficult even for human analysts to interpret consistently, let alone AI models. This paper presents findings from the Santa Clara Microsoft Practicum Project, led by Professor Charlie Goldenberg, which benchmarks the performance of Microsoft's Copilot, OpenAI's ChatGPT, Google's Gemini, and traditional machine learning models for sentiment analysis of financial text. Using Microsoft earnings call transcripts, the analysis assesses how well LLM-derived sentiment correlates with market sentiment and stock movements and evaluates the accuracy of model outputs. Prompt engineering techniques are also examined to improve sentiment analysis results. Visualizations of sentiment consistency are developed to evaluate alignment between tone and stock performance, with sentiment trends analyzed across Microsoft's lines of business to determine which segments exert the greatest influence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。