测试大模型金融智商,发现超级投资AI表现最佳。
Evaluating Financial Intelligence in Large Language Models: Benchmarking SuperInvesting AI with LLM Engines
- 构建五维评估框架,测大模型金融分析能力。
- 超级投资AI准确率达8.96/10,完成度56.65/70最高。
- 结合数据与推理的模型更适合复杂投研任务。
大型语言模型在金融分析和投资研究中应用日益广泛,但其金融推理能力的系统性评估仍显不足。本文提出AI金融智能基准(AFIB),一个涵盖事实准确性、分析完整性、数据时效性、模型一致性及失败模式五个维度的多维评估框架。我们使用来自真实股票研究任务的95+个结构化金融分析问题,对GPT、Gemini、Perplexity、Claude和SuperInvesting五个AI系统进行评估。结果表明各模型表现差异显著:SuperInvesting在整体性能上最优,平均事实准确率为8.96/10,分析完整性得分56.65/70,且幻觉率最低;检索型系统如Perplexity因可访问实时信息,在数据时效性任务上表现优异,但在分析综合与一致性方面较弱。总体表明,金融智能是多维的,兼具结构化数据获取与分析推理能力的系统在复杂投资研究中表现最可靠。
原文摘要 · Abstract (English)
Large language models are increasingly used for financial analysis and investment research, yet systematic evaluation of their financial reasoning capabilities remains limited. In this work, we introduce the AI Financial Intelligence Benchmark (AFIB), a multi-dimensional evaluation framework designed to assess financial analysis capabilities across five dimensions: factual accuracy, analytical completeness, data recency, model consistency, and failure patterns. We evaluate five AI systems: GPT, Gemini, Perplexity, Claude, and SuperInvesting, using a dataset of 95+ structured financial analysis questions derived from real-world equity research tasks. The results reveal substantial differences in performance across models. Within this benchmark setting, SuperInvesting achieves the highest aggregate performance, with an average factual accuracy score of 8.96/10 and the highest completeness score of 56.65/70, while also demonstrating the lowest hallucination rate among evaluated systems. Retrieval-oriented systems such as Perplexity perform strongly on data recency tasks due to live information access but exhibit weaker analytical synthesis and consistency. Overall, the results highlight that financial intelligence in large language models is inherently multi-dimensional, and systems that combine structured financial data access with analytical reasoning capabilities provide the most reliable performance for complex investment research workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。