评测大模型在金融场景下的发散与收敛思维能力
Reasoning Beyond the Obvious: Evaluating Divergent and Convergent Thinking in LLMs for Financial Scenarios
- 构建双维度金融推理基准ConDiFi,涵盖发散与收敛任务
- GPT-4o虽流畅但创新性与可操作性不足,深求、Cohere表现更优
- 适合关注AI投资决策能力的研究者与从业者
现有大模型推理评测多聚焦事实准确性或步骤逻辑,但在金融领域,专业人士不仅需得出最优结论,还需在不确定性下提出富有创意且合理的未来情景。本文提出ConDiFi基准,同时评估大模型在金融任务中的发散与收敛思维能力。该基准包含607个宏观金融场景的发散式推理题和990道多跳对抗性选择题用于收敛推理评估。基于此,我们评测了14个主流模型,发现尽管GPT-4o在语言流畅性上表现优异,但在新颖性(Novelty)与可操作性(Actionability)方面表现较弱;而DeepSeek-R1与Cohere Command R+在生成可指导投资的洞察方面位居前列。ConDiFi为评估大模型在金融中安全、战略性部署所需的核心推理能力提供了新视角。
原文摘要 · Abstract (English)
Most reasoning benchmarks for LLMs emphasize factual accuracy or step-by-step logic. In finance, however, professionals must not only converge on optimal decisions but also generate creative, plausible futures under uncertainty. We introduce ConDiFi, a benchmark that jointly evaluates divergent and convergent thinking in LLMs for financial tasks. ConDiFi features 607 macro-financial prompts for divergent reasoning and 990 multi-hop adversarial MCQs for convergent reasoning. Using this benchmark, we evaluated 14 leading models and uncovered striking differences. Despite high fluency, GPT-4o underperforms on Novelty and Actionability. In contrast, models like DeepSeek-R1 and Cohere Command R+ rank among the top for generating actionable, insights suitable for investment decisions. ConDiFi provides a new perspective to assess reasoning capabilities essential to safe and strategic deployment of LLMs in finance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。