大模型金融情感分析中,直觉判断比深度推理更准确。
Reasoning or Overthinking: Evaluating Large Language Models on Financial Sentiment Analysis
- 用系统1/系统2思维模拟对比模型表现,发现无需推理提示
- GPT-4o无CoT提示时最接近人工标注结果
- 复杂语言和标注分歧会引发过度思考,降低准确性
我们研究了大语言模型(包括基于推理与非推理的模型)在零样本金融情感分析中的有效性。基于领域专家标注的Financial PhraseBank数据集,评估不同LLM及提示策略在金融语境下与人工标注情感的一致性。比较三种专有模型(GPT-4o、GPT-4.1、o3-mini)在模拟系统1(快速直觉)与系统2(缓慢推理)思维的不同提示范式下的表现,并与两个在金融情感分析上微调的小模型(FinBERT-Prosus、FinBERT-Tone)进行基准对比。结果表明,推理(无论是通过提示还是模型设计)并未提升该任务表现。令人意外的是,最准确且与人类一致的组合是未使用链式思维(CoT)提示的GPT-4o。我们进一步分析语言复杂度与标注一致性对性能的影响,发现推理可能引发过度思考,导致预测不佳。这表明,在金融情感分类任务中,类似系统1的快速直觉判断比系统2式的慢速推理更贴近人类判断。研究挑战了‘更多推理=更好决策’的默认假设,尤其在高风险金融应用中。
原文摘要 · Abstract (English)
We investigate the effectiveness of large language models (LLMs), including reasoning-based and non-reasoning models, in performing zero-shot financial sentiment analysis. Using the Financial PhraseBank dataset annotated by domain experts, we evaluate how various LLMs and prompting strategies align with human-labeled sentiment in a financial context. We compare three proprietary LLMs (GPT-4o, GPT-4.1, o3-mini) under different prompting paradigms that simulate System 1 (fast and intuitive) or System 2 (slow and deliberate) thinking and benchmark them against two smaller models (FinBERT-Prosus, FinBERT-Tone) fine-tuned on financial sentiment analysis. Our findings suggest that reasoning, either through prompting or inherent model design, does not improve performance on this task. Surprisingly, the most accurate and human-aligned combination of model and method was GPT-4o without any Chain-of-Thought (CoT) prompting. We further explore how performance is impacted by linguistic complexity and annotation agreement levels, uncovering that reasoning may introduce overthinking, leading to suboptimal predictions. This suggests that for financial sentiment classification, fast, intuitive "System 1"-like thinking aligns more closely with human judgment compared to "System 2"-style slower, deliberative reasoning simulated by reasoning models or CoT prompting. Our results challenge the default assumption that more reasoning always leads to better LLM decisions, particularly in high-stakes financial applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。