arXiv:2503.18313cs.MAcs.AI2025-03被引 4

用真实市场环境测试大模型炒股能力,发现现有评测有漏洞。

Will LLMs be Professional at Fund Investment? DeepFund: A Live Arena Perspective

  • 构建多智能体系统模拟真实投资流程
  • 在真实市场中验证模型表现,避免数据泄露问题
  • 适合研究AI金融应用或评估大模型实战能力者

大语言模型在多个领域表现出色,但在金融决策中的实际效果仍缺乏充分评估。现有基准主要测试模型对金融文档的理解,而非在动态市场中管理资产或挖掘交易机会的能力。尽管已有新基准覆盖金融领域的多样化任务,我们识别出四大缺陷:数据泄露、自我循环、过度干预和维护困难。为此,我们提出DeepFund——一个用于评估基于LLM的交易策略的实时竞技平台。该平台采用多智能体框架,模拟现实中投资决策的关键角色,并提供可视化网页界面,展示模型在不同市场条件下的基金投资指标,支持细致比较分析。通过DeepFund,旨在为大模型在基金投资中的能力提供更真实、公平的评估,揭示其在真实金融市场的潜力。代码已开源:https://github.com/HKUSTDial/DeepFund。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated impressive capabilities across various domains, but their effectiveness in financial decision-making remains inadequately evaluated. Current benchmarks primarily assess LLMs' understanding on financial documents rather than the ability to manage assets or dig out trading opportunities in dynamic market conditions. Despite the release of new benchmarks for evaluating diversified tasks on the financial domain, we identified four major problems in these benchmarks, which are data leakage, navel-gazing, over-intervention, and maintenance-hard. To pave the research gap, we introduce DeepFund, a comprehensive arena platform for evaluating LLM-based trading strategies in a live environment. Our approach implements a multi-agent framework where they serve as multiple key roles that realize the real-world investment decision processes. Moreover, we provide a web interface that visualizes LLMs' performance with fund investment metrics across different market conditions, enabling detailed comparative analysis. Through DeepFund, we aim to provide a more realistic and fair assessment on LLM's capabilities in fund investment, offering diversified insights and revealing their potential applications in real-world financial markets. Our code is publicly available at https://github.com/HKUSTDial/DeepFund.

大模型金融投资多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。